archaeobytology.org / textbook / Continuous Reader
Back to Portal
Continuous Reader Mode

Archaeobytology: The Complete Textbook

All 18 chapters, 5 appendices, bibliography, and subject index compiled into a single scrollable manuscript. Use Ctrl+F to search the complete volume.

By Josie Jefferson & Felix Velasco View Curricular Overview
Part I • Theoretical Foundations & Taxonomy

Chapter 1: Introduction to Archaeobytology

19 min read 4,093 words

Opening Vignette: The Day GeoCities Died §

On October 26, 2009, Yahoo announced that GeoCities—one of the earliest and largest web hosting services—would shut down in less than three weeks. What followed was a frantic rescue operation. A loose network of digital preservationists calling themselves “Archive Team” mobilized immediately, recruiting volunteers to scrape as many sites as possible before the November deadline. Working around the clock across time zones, they managed to save approximately 650 gigabytes of data—a fraction of the estimated terabytes that had accumulated since GeoCities launched in 1994.

When the servers went dark on October 26, 2009, roughly 30 million websites vanished from the internet. Personal homepages that teenagers had built in the late 1990s. Fan sites dedicated to Sailor Moon and Dragon Ball Z. Memorial pages for loved ones. Experimental art projects. Tutorials on HTML and web design. An entire era of digital culture—messy, earnest, strange, and deeply human—was murdered by a single corporate decision.

This wasn’t obsolescence. These sites didn’t decay naturally or become technically incompatible. They were deliberately killed by a platform that no longer found them profitable. And because GeoCities users didn’t own their digital ground—they were renting space at geocities.com/neighborhood/username—they had no recourse. When Yahoo turned off the servers, those 30 million voices simply… disappeared.

Or almost disappeared. Thanks to Archive Team’s heroic effort, fragments survived. But even those rescued artifacts present new problems: they exist in a massive, unsearchable data dump. No context. No curation. No way to understand why a particular site mattered or what community it represented. The bits are preserved, but the meaning is lost.

This is the problem Archaeobytology was created to solve.


What Is Archaeobytology? §

Archaeobytology is the study and practice of excavating, preserving, interpreting, and building with digital artifacts—particularly those that have been “murdered” by platform shutdowns or rendered obsolete by technological change.

The term combines three roots:

  • Archaeo- (Greek: ancient, old) — invoking archaeology’s commitment to recovering and interpreting the past
  • -byte- (computing unit) — grounding the discipline in digital materiality
  • -ology (study of) — claiming status as a rigorous intellectual discipline

But Archaeobytology is more than just “digital archaeology.” It encompasses:

  1. Theoretical frameworks for understanding digital mortality, platform power, and technological sovereignty
  2. Practical methods for excavation, preservation, and forensic analysis
  3. Institutional design for building organizations that can sustain preservation work for decades
  4. Political advocacy for laws and policies that protect digital culture from corporate erasure
  5. Creative practice of forging new tools, monuments, and systems that embody principles of digital sovereignty

Unlike adjacent fields—digital history, media archaeology, library science, or computer science—Archaeobytology insists on a dual commitment: we are both archivists and builders, both scholars and smiths. We preserve what platforms murder, and we forge alternatives that resist future murders.


Why We Need a New Discipline §

The Inadequacy of Existing Fields §

When GeoCities died, where were the experts? Digital historians could analyze what was lost, but they lacked the technical skills to mount a rescue. Computer scientists could write scraping scripts, but they lacked frameworks for ethical triage or cultural curation. Librarians understood preservation, but most weren’t equipped to reverse-engineer dying platforms or navigate copyright gray areas.

The problem wasn’t lack of expertise—it was fragmentation. The skills needed to address platform death are scattered across multiple disciplines, none of which fully claim this territory:

  • Computer Science treats digital objects as technical problems, not cultural artifacts
  • History studies the past but rarely intervenes to preserve the present
  • Library Science excels at cataloging but often lacks technical depth for complex digital formats
  • Media Archaeology theorizes dead media but doesn’t always prioritize active preservation
  • Cultural Studies critiques platforms but rarely builds alternatives

Archaeobytology argues that digital preservation requires a unified discipline that combines:

  • Technical skill (excavation, forensics, emulation)
  • Humanistic interpretation (curation, contextualization, ethics)
  • Institutional expertise (sustainable organizations, governance models)
  • Political engagement (policy advocacy, movement building)

No existing field does all of this. That’s why we need a new one.

The Scale of the Crisis §

Platform death isn’t a niche problem. It’s accelerating:

Major Platform Shutdowns (2000-2024):

  • GeoCities (2009): 30 million sites
  • Google Reader (2013): Millions of curated RSS feeds
  • Vine (2017): 200 million short videos
  • Google+ (2019): User profiles, communities, content
  • Mixer (2020): Game streaming platform
  • Tumblr NSFW purge (2018): Millions of posts deleted
  • Twitter’s chaos (2022-present): Uncertain future, mass exodus
  • Countless smaller platforms: Ello, Peach, Path, Plurk, FriendFeed…

Each shutdown represents not just technical infrastructure dying, but communities, memories, identities, and cultural artifacts being erased. And unlike physical artifacts—which decay slowly, giving civilizations time to respond—digital artifacts can vanish overnight.

Moreover, we’re not just losing content. We’re losing:

  • Social graphs: Who was connected to whom
  • Affordances: How platforms shaped communication (Twitter’s 140 characters, Vine’s 6 seconds)
  • Aesthetics: Platform-specific design cultures (GeoCities’ tags, Tumblr’s GIF aesthetic)
  • Communities: Specialized subcultures that formed around platform features

This isn’t just cultural loss—it’s cultural murder. And it’s happening faster than any single discipline can address.


The Three Pillars: A Framework for Digital Sovereignty §

At the heart of Archaeobytology lies a normative commitment: we believe digital culture should be sovereign—independent from corporate control, resistant to platform shutdown, and owned by the people who create it.

We call this framework The Three Pillars, drawing on both ancient philosophy (the Greek concept of authentikos, self-originating authority) and practical infrastructure design. A digitally sovereign presence requires three interdependent foundations:

Pillar 1: Declaration (I Am) §

The Principle: You should be able to declare your identity and existence without permission from a platform or intermediary.

When you create a Facebook profile, Facebook owns your identity. If they ban you, “you” cease to exist—at least in that digital space. Your name, photos, relationships, and history are controlled by an entity that can revoke them at will.

Sovereign declaration means:

  • Your identity is self-hosted ([email protected], not [email protected])
  • Your presence is persistent (your website outlives any platform)
  • Your voice is uncensorable (no corporate ToS can silence you on your own ground)

This doesn’t mean freedom from all consequences—legal systems, community norms, and social accountability still apply. But it means platforms cannot unilaterally erase you.

Historical precedent: In the early web (1990s-2000s), personal homepages embodied this principle. You bought a domain, hosted your site, and declared yourself to the internet. GeoCities degraded this model by making addresses hierarchical (geocities.com/neighborhood/you) rather than sovereign (you.com). Social media completed the enclosure by eliminating personal domains entirely.

Pillar 2: Connection (Instant Message) §

The Principle: You should be able to communicate directly with others without a platform mediating, monitoring, or monetizing your relationships.

Platforms don’t just host our content—they control our connections. Facebook decides who sees your posts (algorithmic curation). Twitter can prevent you from messaging someone (shadowbanning). Instagram owns the graph of your followers (you can’t export it).

Sovereign connection means:

  • Communication is peer-to-peer or federated (not routed through corporate servers)
  • Relationships are exportable (if you leave a platform, your network comes with you)
  • Discovery is intentional (you choose who to connect with, not an algorithm)

This is why email—for all its flaws—remains more sovereign than social media. If Gmail shuts down, you can take your address to another provider. If your contacts have their own domains ([email protected]), you can reach them directly.

The challenge: Network effects make this hard. If everyone is on Twitter, leaving Twitter means losing access to your community. Sovereignty requires interoperability—the ability to communicate across platforms, or to bring your network with you when you migrate.

Pillar 3: Ground (Digital Real Estate) §

The Principle: You should own the infrastructure your digital life is built on—not rent it from a landlord who can evict you.

GeoCities users thought they had websites. They didn’t. They had leases on someone else’s servers. When Yahoo decided those leases weren’t profitable, they terminated them. No appeals, no alternatives, no recourse.

Sovereign ground means:

  • You own your domain name (example.com, not facebook.com/example)
  • You control your hosting (self-hosted or a provider you can migrate away from)
  • Your data is exportable (you can take it with you, in usable formats)

This doesn’t require technical expertise. Thousands of people own domains and use managed hosting services like WordPress.com or Ghost(Pro). The key is portability: if the service shuts down or changes terms, you can move.

The analogy: Owning ground is like owning land versus renting an apartment. A landlord can raise rent, change rules, or evict you. But if you own land, you have sovereignty—subject to laws, but not to arbitrary corporate power.


The Dual Soul: Archive and Anvil §

Archaeobytology is not just a preservationist discipline. We insist on a dual practice:

The Archive: Preservation and Memory §

The Archive represents our commitment to:

  • Excavate murdered platforms before they vanish completely
  • Preserve artifacts with technical and cultural fidelity
  • Curate collections that make sense of vast data dumps
  • Interpret artifacts so future generations understand their significance

Archival work requires:

  • Technical skills (web scraping, forensic recovery, emulation)
  • Ethical frameworks (what should be preserved? what should be forgotten?)
  • Institutional knowledge (how do you build organizations that last 50 years?)

The Archive is retrospective: it looks backward to save what’s endangered.

The Anvil: Creation and Resistance §

The Anvil represents our commitment to:

  • Forge tools that empower digital sovereignty (domain registration services, self-hosting platforms, open protocols)
  • Build monuments that embody our values (websites, platforms, frameworks designed for permanence)
  • Design institutions that resist platform capture (non-profit archives, cooperative hosting services, federated networks)

The work of the Anvil requires:

  • Creative practice (making things that don’t yet exist)
  • Systems thinking (how do you design for resilience?)
  • Political imagination (what does a post-platform future look like?)

The Anvil is prospective: it looks forward to build alternatives.

Why Both Are Necessary §

You cannot be only an archivist. If you preserve everything but build nothing, you’re a curator in a warehouse—keeping records of a world dominated by platforms, never challenging that dominance.

You cannot be only a builder. If you forge alternatives but never preserve the past, you lose the lessons of history. Each new generation reinvents the wheel, repeating old mistakes.

The Archaeobytologist embodies both: we save the murdered web, and we build systems that can’t be murdered.


Triage: The Central Methodology §

The most painful truth of Archaeobytology: you cannot save everything.

When a platform announces shutdown, you have limited time, limited storage, limited volunteers. You must make choices. This is triage—borrowed from emergency medicine, where doctors must decide which patients to treat first when resources are scarce.

The Custodial Filter §

We use the Custodial Filter as our ethical framework for triage decisions. Before preserving an artifact, we ask five questions:

  1. Cultural Significance: Does this artifact represent a community, movement, or cultural moment that would otherwise be lost?

  2. Technical Fragility: How close to disappearance is this? (A site archived by Internet Archive is less urgent than one that isn’t.)

  3. Rescue Difficulty: How hard is this to preserve? (Simple HTML is easier than complex Flash applications.)

  4. Existing Redundancy: Is someone else already preserving this? (Don’t duplicate effort when time is scarce.)

  5. Consent and Ethics: Should we preserve this? Does it violate someone’s privacy, contain traumatic content, or cause harm by existing?

The fifth question is critical. Not everything that can be preserved should be. Revenge porn, doxxing, harassment campaigns—these are digital artifacts too, but preserving them can perpetuate harm. The Custodial Filter requires us to think beyond technical feasibility to ethical responsibility.

Triage in Practice: The Archive Team Model §

When Vine announced its shutdown in 2016, Archive Team had roughly six weeks to save 200 million videos. Impossible to save them all. They triaged:

  • Highest priority: Videos with significant cultural impact (viral memes, influential creators, historically important moments)
  • Medium priority: Representative samples of different communities, genres, and time periods
  • Lower priority: Duplicates, spam, commercial advertisements

Even then, they couldn’t manually curate 200 million items. So they used algorithmic triage: view counts, shares, and community-submitted nominations. Imperfect, but pragmatic.

The result: they saved millions of videos, but not all. Some Vines are lost forever. Triage accepts this tragedy as unavoidable, while working to minimize the loss.


A Brief History of Digital Mortality §

Digital culture has always been ephemeral, but the causes of mortality have evolved:

Era 1: Technological Obsolescence (1960s-1990s) §

Early digital artifacts died because the hardware or software became incompatible:

  • Floppy disks degraded physically
  • File formats became unreadable (WordStar, AppleWorks)
  • Storage media evolved (punch cards → magnetic tape → hard drives)

This was passive death—artifacts decayed like ancient papyrus. The solution was technical: emulation, format migration, hardware preservation.

As the web grew, artifacts died because:

  • Website owners stopped paying for hosting
  • Domains expired and were re-registered by squatters
  • Links broke as sites moved or vanished

This was death by neglect—the equivalent of abandoning a physical archive to water damage and mold. The solution was institutional: projects like the Internet Archive’s Wayback Machine, which proactively crawled and preserved sites.

Era 3: Platform Murder (2000s-present) §

In the social media era, artifacts die because platforms choose to kill them:

  • Corporate shutdowns (GeoCities, Vine, Google Reader)
  • Terms of Service purges (Tumblr NSFW ban, YouTube’s algorithmic demonetization)
  • Acquisition and closure (platforms bought and killed by competitors)

This is active murder—deliberate erasure. The solution isn’t just technical or institutional—it’s political. We need laws, rights, and alternatives.

Archaeobytology emerged in response to this third era. We’re not just fighting entropy or neglect—we’re fighting corporate power.


What Makes Archaeobytology Different? §

Not Digital History §

Digital historians study the past. Archaeobytologists intervene in the present to create a future past. When Archive Team scraped GeoCities, they weren’t analyzing history—they were making it possible for future historians to have something to analyze.

Not Media Archaeology §

Media archaeologists theorize dead media. Archaeobytologists rescue dying media before they become dead. We’re applied, not purely theoretical. We get our hands dirty with code, servers, and scrapers.

Not Library Science §

Librarians excel at cataloging, access, and preservation—but within established frameworks. Archaeobytology operates in legal and technical gray areas: scraping platforms that didn’t consent, preserving copyrighted material under dubious fair use claims, reverse-engineering proprietary formats.

We respect librarians deeply. But we do things they often can’t or won’t do.

Not Computer Science §

Computer scientists can write scrapers and build emulators. But they often lack frameworks for curation, ethics, and cultural interpretation. A computer scientist might preserve every byte. An Archaeobytologist asks: Should we? What does this mean? How do we make it legible?

Not Just Activism §

Archaeobytology isn’t pure advocacy. We build theoretical frameworks, develop rigorous methods, and create institutions. We’re scholars and activists—but the scholarship matters.


The Crisis of Legitimacy §

Archaeobytology faces a credibility problem: we don’t exist yet.

There are no Archaeobytology departments at universities. No tenure-track jobs with “Archaeobytologist” in the title. No dedicated funding streams from NSF or NEH. When we tell people we’re Archaeobytologists, they ask, “What’s that?”

This book is part of solving that problem. By codifying our theories, methods, and practices, we make the discipline real. By teaching courses, publishing research, and building institutions, we establish legitimacy.

Disciplines don’t emerge naturally—they’re constructed through collective action:

  • Journals and conferences create scholarly community
  • Textbooks standardize knowledge
  • Degree programs train new generations
  • Professional organizations provide structure
  • Public advocacy wins recognition

This book is a founding document. You’re reading it early in the discipline’s life. In 20 years, Archaeobytology might be as established as Data Science or Digital Humanities. Or it might remain a niche practice, known only to specialists.

That outcome depends on us.


Who Is This Book For? §

Undergraduate Students §

If you’re considering a career in digital preservation, museum curation, or tech ethics, this book provides foundational knowledge. Each chapter includes exercises and case studies to build practical skills.

Graduate Students and Researchers §

If you’re writing a dissertation on platform death, digital memory, or technological sovereignty, this book offers theoretical frameworks and methodologies you can adapt.

Practitioners §

If you work in libraries, archives, museums, or tech companies, this book gives you tools to advocate for preservation work and design sustainable institutions.

Activists and Advocates §

If you’re fighting for digital rights, platform accountability, or data sovereignty, this book provides evidence and arguments for policy change.

The Curious Public §

If you’ve ever wondered what happened to your MySpace profile, your LiveJournal, or that website you made in 2003, this book explains why they disappeared—and what we can do about it.


How to Use This Book §

Structure §

The book is organized into five parts:

Part I: Foundations (Chapters 1-6) introduces core concepts: the Archaeobyte taxonomy, the Archive/Anvil framework, the Three Pillars, and triage methodology.

Part II: Excavation & Forensics (Chapters 7-10) teaches practical methods for recovering and analyzing digital artifacts.

Part III: Institution Building (Chapters 11-14) shows how to design organizations that sustain preservation work for decades.

Part IV: Systems & Movements (Chapters 15-16) addresses political economy: who controls digital infrastructure, and how do we build alternatives?

Part V: Public Scholarship & The Future (Chapters 17-18) explores how Archaeobytologists can translate research into public discourse, policy, and cultural change.

Pedagogy §

Each chapter includes:

  • Case Studies: Real-world examples of platform death, preservation projects, and institution-building
  • Discussion Questions: Prompts for classroom or reading group conversations
  • Exercises: Hands-on activities to build skills (accessible to readers with varying technical backgrounds)
  • Further Reading: Curated bibliography for deeper exploration

Teaching with This Book §

For a 15-week undergraduate survey course, cover one chapter per week. Focus on Part I (Foundations) and Part II (Methods), with selected chapters from Part III.

For a graduate seminar, assume students have read the entire book. Use class time for deep discussion of case studies, triage dilemmas, and institutional design challenges. Assign a capstone project: design a preservation organization, memory institution, or movement campaign.

For professional development, organize a reading group among librarians, archivists, or tech workers. Each week, one person presents a chapter and leads discussion. Focus on how frameworks apply to your workplace.


A Provocation: Why Bother? §

Let’s be honest: most people don’t care that GeoCities died. They don’t think about digital preservation. They assume “the internet remembers everything” (it doesn’t) or that “tech companies will handle it” (they won’t).

So why bother? Why build a discipline around saving things most people forgot existed?

Three answers:

1. Memory Is Power §

Who controls the past controls the present. Platforms curate our memories—deciding which photos Facebook shows you in “On This Day,” which tweets trend, which YouTube videos get recommended. When platforms die, they take our memories with them.

Preserving murdered platforms is an act of resistance against corporate memory control. It insists that our digital lives belong to us, not to companies that can erase them at will.

2. Culture Dies in Darkness §

Every generation deserves access to the cultural artifacts of previous generations. Historians study ancient Rome through pottery fragments. Future historians will study early internet culture through GeoCities sites—if we save them.

If we don’t preserve digital culture, it vanishes. No ruins, no fragments. Just absence. Future generations won’t even know what they’re missing.

3. Building Alternatives Requires Understanding Failures §

You can’t design a sovereign internet if you don’t understand how platforms murdered the old one. Every shutdown teaches lessons:

  • Why didn’t GeoCities users own their domains?
  • Why couldn’t Vine videos be exported?
  • Why did Google Reader’s death destroy thousands of curated feeds?

Studying murdered platforms isn’t nostalgia—it’s learning how to build systems that can’t be murdered.


The Archaeobytologist’s Vow §

As you read this book, you’re joining a community. We’re small now—scattered practitioners, archivists, activists, scholars. But we’re growing.

If you embrace this work, you’re making an implicit commitment:

I will not let digital culture die in silence.

I will excavate what platforms murder.

I will build systems that resist future murder.

I will teach others to do the same.

I am a scholar and a smith, a custodian and a strategist.

I own my ground. I tell my story. I forge my future.

I am an Archaeobytologist.


Looking Ahead §

The rest of this book will equip you with:

  • Theory: Rigorous frameworks for understanding digital death (Chapters 2-6)
  • Methods: Practical skills for excavation and preservation (Chapters 7-10)
  • Institutions: Models for sustainable organizations (Chapters 11-14)
  • Systems: Alternatives to platform capitalism (Chapters 15-16)
  • Impact: Strategies for public scholarship and policy change (Chapters 17-18)

By the end, you’ll be able to:

  • Classify digital artifacts using the Archaeobyte taxonomy
  • Conduct triage using the Custodial Filter
  • Excavate a dying platform before it shuts down
  • Design a preservation organization that can last 50 years
  • Advocate for laws that protect digital culture
  • Build alternatives that embody digital sovereignty

You’ll be, in short, an Archaeobytologist.

Welcome to the discipline. Now let’s get to work.


Discussion Questions §

  1. On GeoCities: Why do you think Yahoo shut down GeoCities instead of maintaining it as a historical archive? What does this decision reveal about corporate priorities?

  2. On Definitions: How is Archaeobytology different from “digital archiving” or “data preservation”? Does it need to be a separate discipline, or could existing fields do this work?

  3. On The Three Pillars: Audit your own digital presence. Do you have Declaration (sovereign identity)? Connection (direct communication)? Ground (owned infrastructure)? If not, what would it take to achieve them?

  4. On Triage: Imagine a platform announces shutdown in 48 hours. You can save 10% of its content. How do you decide what to save? What ethical dilemmas arise?

  5. On The Dual Soul: Can you be only an archivist (save the past) without being a builder (create the future)? Or are both commitments necessary?

  6. On Legitimacy: What would it take for Archaeobytology to be recognized as a legitimate academic discipline? Journals? Conferences? University departments? All of the above?


Exercise: Your First Triage §

Scenario: You discover that Ello (a social network launched in 2014 as an “ad-free alternative” to Facebook) is shutting down in one week. You have time to preserve approximately 1,000 user profiles out of 50,000 active accounts.

Task:

  1. Research Ello: What communities formed there? What made it culturally significant?
  2. Define Criteria: Using the Custodial Filter, list 5 criteria you’d use to select profiles
  3. Identify Examples: Find 10 specific Ello users you’d prioritize and explain why
  4. Ethical Dilemmas: Identify at least 3 ethical challenges in this scenario (privacy, consent, harm, etc.)
  5. Reflection: After making your choices, what did you have to leave behind? How does that feel?

Further Reading §

Foundational Texts §

  • Kirschenbaum, Matthew. Mechanisms: New Media and the Forensic Imagination. MIT Press, 2008.
  • Parikka, Jussi. What Is Media Archaeology? Polity, 2012.
  • Chun, Wendy Hui Kyong. “The Enduring Ephemeral, or the Future Is a Memory.” Critical Inquiry 35, no. 1 (2008): 148-171.

On Platform Death §

  • Gillespie, Tarleton. Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions That Shape Social Media. Yale University Press, 2018.
  • Zuboff, Shoshana. The Age of Surveillance Capitalism. PublicAffairs, 2019.
  • Doctorow, Cory. The Internet Con: How to Seize the Means of Computation. Verso, 2023.

On Digital Sovereignty §

  • Lessig, Lawrence. Code: Version 2.0. Basic Books, 2006.
  • Schneier, Bruce. Data and Goliath: The Hidden Battles to Collect Your Data and Control Your World. W.W. Norton, 2015.
  • Véliz, Carissa. Privacy Is Power: Why and How You Should Take Back Control of Your Data. Melville House, 2020.

On Archives and Memory §

  • Derrida, Jacques. Archive Fever: A Freudian Impression. University of Chicago Press, 1996.
  • Ernst, Wolfgang. Digital Memory and the Archive. University of Minnesota Press, 2013.
  • Manoff, Marlene. “Theories of the Archive from Across the Disciplines.” Portal: Libraries and the Academy 4, no. 1 (2004): 9-25.

Primary Sources §

  • Archive Team. “GeoCities: We Didn’t Start the Fire.” https://archiveteam.org/index.php?title=GeoCities
  • Internet Archive. Wayback Machine. https://web.archive.org
  • Kahle, Brewster. “Preserving the Internet.” Scientific American 276, no. 3 (1997): 82-83.

End of Chapter 1

Next: Chapter 2 — The Archaeobyte Taxonomy: Understanding Digital Mortality

Part I • Theoretical Foundations & Taxonomy

Chapter 2: The Archaeobyte Taxonomy

Understanding Digital Mortality

23 min read 5,006 words

Opening: Four Artifacts, Four Fates §

Consider four digital objects, each with a different relationship to death:

Artifact 1: A GeoCities homepage from 1998, hosted at geocities.com/SiliconValley/1234. When Yahoo shut down GeoCities in 2009, this site vanished from the live web. But Archive Team scraped it before shutdown, and it now exists as files in a 650GB torrent. The site is dead but preserved—murdered by its platform, rescued by volunteers, waiting for someone to resurrect it.

Artifact 2: Your current Twitter profile, actively maintained with daily posts. But Twitter’s future is uncertain. Elon Musk’s chaotic ownership has driven mass exodus to alternatives. Your profile is alive but endangered—functioning now, but vulnerable to corporate whims, algorithmic changes, or eventual shutdown.

Artifact 3: A Flash game called Homestar Runner, created in the early 2000s. Flash Player was discontinued by Adobe in 2020, rendering millions of Flash games unplayable. The files still exist, but without emulation software, they’re inert. This artifact is technically dead but spiritually haunting—the data persists, but the experience is inaccessible without intervention.

Artifact 4: An ancient stone tablet with cuneiform writing, created 4,000 years ago in Mesopotamia. The civilization that made it is long gone, but the tablet survives in a museum. It’s long-dead but monumentally preserved—so old that its mortality is complete, yet so durably encoded that it outlasted empires.

These four artifacts represent four fundamentally different states of digital mortality. Archaeobytology needs a taxonomy to distinguish between them—not just for academic precision, but for practical triage. Each state demands different preservation strategies, ethical considerations, and urgency levels.

This chapter introduces the Archaeobyte Taxonomy: a classification system for understanding how digital artifacts live, die, and persist.


The Archaeobyte Taxonomy: Four Categories §

We classify digital artifacts into four types based on their mortality state and preservation status:

1. Archaeobyte (Dead and Preserved) §

2. Vivibyte (Alive and Endangered) §

3. Umbrabyte (Dead but Haunting) §

4. Petribyte (Monumentally Preserved) §

Each category has distinct characteristics, ethical challenges, and preservation needs. Let’s examine them in depth.


1. Archaeobyte: The Dead and Preserved §

Definition §

An Archaeobyte is a digital artifact that:

  • Was once alive (accessible, functional, part of active culture)
  • Died through platform shutdown, obsolescence, or deliberate deletion
  • Has been preserved in some form (archived, scraped, backed up)
  • Exists in a liminal state between death and potential resurrection

The term combines archaeo- (ancient, belonging to the past) with byte (digital information unit). Archaeobytes are digital fossils—remnants of dead platforms waiting to be excavated, interpreted, and potentially revived.

Characteristics §

Temporality: Archaeobytes occupy past time. They were created years or decades ago and reflect the technological, cultural, and social contexts of their era.

Accessibility: They exist in archives but aren’t easily accessible. You can’t just visit a URL. You need to know where the archive is, how to navigate it, and potentially how to run emulation software.

Functionality: Some Archaeobytes are fully functional if properly emulated (a Flash game can still be played). Others are fragmentary (HTML pages with broken images, databases without their front-ends).

Cultural Context: Archaeobytes often lack context. A GeoCities page preserved as raw HTML doesn’t tell you who made it, why it mattered, or what community it belonged to. Interpretation requires detective work.

Examples §

GeoCities Archives (Archive Team, 2009)

  • 650GB torrent containing millions of HTML files
  • Raw dump with minimal metadata
  • Requires local server to view properly
  • Missing images, broken links, no search functionality
  • Status: Preserved but not curated

Vine Archive (Internet Archive, 2017)

  • Millions of 6-second videos scraped before shutdown
  • Stored in Internet Archive’s video collection
  • Searchable by creator username
  • Videos playable but divorced from original social context (comments, likes, loops)
  • Status: Preserved with partial metadata

Flash Games (Flashpoint Project, ongoing)

  • 500,000+ Flash games and animations preserved
  • Requires custom launcher with embedded emulator
  • Fully playable with original functionality
  • Community-curated with descriptions and tags
  • Status: Preserved, curated, and resurrected

CD-ROM Multimedia (Internet Archive, various dates)

  • 1990s educational software, encyclopedias, games
  • Runs in browser via emulation
  • Often includes full documentation and original packaging scans
  • Status: Museum-quality preservation

Preservation Needs §

Archaeobytes require:

  1. Storage infrastructure: Servers, hard drives, distributed backups
  2. Emulation or compatibility layers: Flash emulators, old browser engines, virtual machines
  3. Metadata and contextualization: Who made this? When? Why did it matter?
  4. Access systems: Search, browse, discovery mechanisms
  5. Legal frameworks: Copyright exceptions for preservation (often operating in gray areas)

Ethical Considerations §

Consent: Did the original creators consent to preservation? Many GeoCities users abandoned their sites and might not want them resurrected.

Privacy: Personal information (emails, addresses, photos) embedded in old sites may violate current privacy expectations.

Context Collapse: Artifacts created for small communities (a private forum, a friend’s homepage) now exposed to anyone who finds the archive.

Authenticity: When you emulate a Flash game, is it still the “same” artifact? Or is emulation a form of transformation?

Triage Priority: Medium §

Archaeobytes are already preserved, so they’re not in immediate danger of disappearing entirely. But they’re at risk of:

  • Bit rot: Storage media degrading over time
  • Format obsolescence: Emulators becoming outdated
  • Link rot: Archives moving or disappearing
  • Institutional failure: Organizations running archives may shut down

Priority increases if:

  • No redundant copies exist elsewhere
  • The archive is hosted by a fragile organization
  • The artifacts have high cultural significance

2. Vivibyte: The Alive and Endangered §

Definition §

A Vivibyte is a digital artifact that:

  • Is currently alive (accessible, functional, actively used)
  • Exists on vulnerable infrastructure (commercial platforms, centralized servers, proprietary systems)
  • Faces existential threats (platform instability, corporate acquisition, terms of service changes, economic precarity)

The term combines vivi- (living, alive) with byte. Vivibytes are the living endangered species of digital culture—thriving now but facing extinction.

Characteristics §

Temporality: Vivibytes are present-tense. They’re being created, updated, and used right now.

Accessibility: They’re easily accessible—just visit a URL. But that accessibility is contingent on platform stability.

Dependency: Vivibytes depend on platform infrastructure. If Twitter shuts down, every tweet becomes inaccessible (unless archived).

Precarity: Their survival isn’t guaranteed. They exist at the mercy of corporate decisions, algorithm changes, and terms of service enforcement.

Examples §

Twitter/X (2023-present)

  • Elon Musk’s acquisition created massive instability
  • Mass layoffs gutted engineering and trust & safety teams
  • API restrictions killed third-party clients
  • Unpredictable policy changes (verified checkmarks, algorithmic timeline changes)
  • Mass user exodus to Mastodon, Bluesky, Threads
  • Status: Alive but facing existential crisis

Substack Newsletters (ongoing)

  • Writers build audiences on Substack’s platform
  • Substack owns the domain (username.substack.com)
  • Export tools exist but are imperfect (subscriber lists can be exported, but URLs break if you move)
  • Vulnerable to Substack’s business model changes
  • Recent controversies over content moderation have driven some writers to Ghost or self-hosted options
  • Status: Alive, functional, but sovereignty questions emerging

TikTok (2020-present)

  • Facing potential US ban due to national security concerns
  • Creators have millions of followers but no platform ownership
  • Videos are proprietary format, difficult to export
  • Algorithm is opaque and changes frequently
  • Status: Thriving but politically endangered

Discord Servers (ongoing)

  • Millions of communities hosted on proprietary platform
  • Chat history owned by Discord, not communities
  • No easy export of full server history
  • Vulnerable to Discord’s moderation policies and business decisions
  • If Discord shuts down, all communities vanish
  • Status: Alive, widely used, but entirely dependent on corporate stability

Indie Web Personal Sites (various)

  • Bloggers using self-hosted WordPress or static site generators
  • Own their domains and content
  • Less vulnerable to platform shutdown
  • But still depend on hosting providers, domain registrars, and web standards
  • Status: More sovereign than platform-hosted content, but not invulnerable

Preservation Needs §

Vivibytes require proactive archiving:

  1. Continuous crawling: Internet Archive’s Wayback Machine constantly archives the live web
  2. User-driven backups: Individuals exporting their own data (Twitter archives, Instagram data downloads)
  3. Institutional partnerships: Libraries and archives working with platforms to preserve content before shutdown
  4. Legal preparation: Advocacy for “right to archive” laws that allow preservation without permission

Ethical Considerations §

Timing: When do you preserve a Vivibyte? If you archive someone’s tweets daily, are you violating their expectation of ephemeral communication?

Comprehensiveness: Should you archive everything on a platform, or only what’s “culturally significant”? Who decides?

Privacy: Many Vivibytes contain personal information shared with the expectation that it will disappear eventually. Permanent archiving changes that expectation.

Platform Relationships: Should archivists work with platforms (negotiated data dumps) or against them (scraping without permission)?

Triage Priority: Variable (Low to Critical) §

Priority depends on threat imminence:

  • Low: Stable platforms with good export tools (WordPress.com, GitHub)
  • Medium: Platforms with uncertain futures but no immediate danger (Reddit, Medium)
  • High: Platforms showing signs of instability (mass layoffs, leadership chaos, user exodus)
  • Critical: Platforms that have announced shutdown (weeks or months remaining)

Indicators of rising threat:

  • Leadership changes or acquisitions
  • Financial struggles (layoffs, failed funding rounds)
  • User exodus or declining engagement
  • Policy changes that anger core communities
  • Technical instability (outages, bugs)
  • Legal or regulatory threats

3. Umbrabyte: The Dead but Haunting §

Definition §

An Umbrabyte is a digital artifact that:

  • Is technically dead (inaccessible, non-functional, or obsolete)
  • Has not been properly preserved (exists in fragmentary or corrupted form)
  • Haunts the present through memory, references, or partial remnants
  • Could theoretically be resurrected with sufficient effort, but currently exists in limbo

The term combines umbra- (shadow, ghost) with byte. Umbrabytes are digital ghosts—artifacts that are neither fully alive nor fully preserved, occupying a haunting middle ground.

Characteristics §

Temporality: Umbrabytes are caught between past and present. They died, but they haven’t been properly mourned or memorialized.

Accessibility: They exist in fragments—dead links, broken images, corrupted files, screenshots, memories.

Liminality: Umbrabytes occupy a liminal state. They’re dead enough to be inaccessible but alive enough to be remembered.

Urgency: Many Umbrabytes are in danger of permanent loss. If not rescued soon, they’ll transition from “dead but haunting” to simply “dead.”

Examples §

MySpace Music (2003-2013)

  • In 2019, MySpace admitted it had “lost” 12 years of user-uploaded music due to a botched server migration
  • Estimated 50 million songs vanished
  • No comprehensive backup exists
  • Some songs survive as MP3s users downloaded
  • Others exist as memories: “I heard this amazing band on MySpace in 2007, but I can’t find them anywhere now”
  • Status: Mostly lost, partially haunting through fragments

Early YouTube (2005-2008)

  • Many early YouTube videos were deleted by users or removed for copyright
  • Internet Archive captured some, but not comprehensively
  • Cultural artifacts like early memes, viral videos, and video responses are often lost
  • Remembered through references, compilations, and oral history
  • Status: Partially preserved, partially lost

Deleted Reddit Communities (various)

  • Reddit has banned thousands of subreddits over the years (r/FatPeopleHate, r/ChapoTrapHouse, r/The_Donald, etc.)
  • Some were archived by volunteers or by Pushshift (academic Reddit archive)
  • Many were not preserved
  • They haunt Reddit culture through references, screenshots, and exile communities that formed elsewhere
  • Status: Partially archived, partially lost, culturally haunting

Flash Websites (1990s-2010s)

  • Millions of Flash-based websites went offline or became non-functional when Flash Player was discontinued in 2020
  • Some are preserved by Flashpoint or Internet Archive
  • Many are lost—known only through screenshots or memories
  • Agency portfolio sites, experimental art projects, interactive storytelling
  • Status: Fragmentarily preserved, largely inaccessible

Private Forums and Message Boards (various)

  • Thousands of small forums shut down over the years (phpBB, vBulletin, etc.)
  • Most weren’t archived by Internet Archive (robots.txt blocks, login walls)
  • Communities lost their entire histories
  • Surviving fragments: Google cache, screenshots, PDFs saved by individual users
  • Status: Largely lost, mourned by former members

Preservation Needs §

Umbrabytes require urgent rescue:

  1. Forensic recovery: Hunting down partial copies, cached pages, user backups
  2. Community archaeology: Interviewing people who remember the artifacts
  3. Reconstruction: Piecing together fragments to create partial records
  4. Metadata creation: Documenting what existed, even if the full artifact can’t be recovered
  5. Triage acceptance: Acknowledging that some Umbrabytes are irrecoverably lost

Ethical Considerations §

Right to Be Forgotten: Some Umbrabytes were intentionally deleted by their creators. Should we resurrect them against their wishes?

Trauma: Some Umbrabytes are traumatic (harassment campaigns, doxxing, revenge porn). Should we let them stay dead?

Reconstructive Violence: Is it ethical to “reconstruct” an artifact from fragments if the result isn’t accurate to the original?

Mourning vs. Resurrection: Sometimes the most ethical response is to mourn an Umbrabyte rather than resurrect it—to acknowledge its loss without trying to recover it.

Triage Priority: Critical (but often futile) §

Umbrabytes are in the most dangerous state:

  • They’re not fully preserved, so they could vanish completely
  • They’re not alive, so there’s no “live source” to capture
  • Time is running out—fragments degrade, memories fade, caches expire

But triage is complicated:

  • Rescue is often technically difficult (fragments scattered, formats corrupted)
  • Success rates are low (many Umbrabytes are irrecoverable)
  • Resources might be better spent on Vivibytes (save the living before mourning the dead)

The hardest triage decisions involve Umbrabytes: Do you spend weeks trying to recover a lost forum’s fragments, or do you focus on archiving a living platform that could die tomorrow?


4. Petribyte: The Monumentally Preserved §

Definition §

A Petribyte is a digital artifact that:

  • Is so old that its original context is historical (decades-old, often pre-web)
  • Has been durably preserved by institutions (libraries, museums, archives)
  • Is treated as cultural heritage (studied by scholars, exhibited in museums)
  • Has achieved stability (no longer at risk of immediate loss)

The term combines petri- (stone, rock—from Latin petra) with byte. Petribytes are digital monuments—artifacts that have achieved the stability of ancient stone tablets, preserved and curated by institutions.

Characteristics §

Temporality: Petribytes are historical. They’re old enough that they’re studied as artifacts of past eras, not current culture.

Accessibility: They’re often highly accessible—digitized, exhibited, documented. Museums and libraries make them available.

Curation: Petribytes receive institutional care—metadata, contextualization, conservation. They’re not just stored; they’re curated.

Monumentality: They’ve achieved cultural recognition. Scholars write about them. Museums exhibit them. They’re canonized.

Examples §

The WELL (1985-present)

  • One of the earliest online communities
  • Archived by Internet Archive and studied by scholars
  • Documented in books like The Virtual Community by Howard Rheingold
  • Still running (as of 2025) but also preserved in multiple forms
  • Status: Monument to early internet culture

Colossal Cave Adventure (1976)

  • Text-based adventure game, one of the first of its kind
  • Source code preserved and studied
  • Multiple versions archived and playable via emulation
  • Influential enough to be analyzed in game studies courses
  • Status: Canon of video game history

ARPANET (1969-1990)

  • Precursor to the internet
  • Decommissioned in 1990, but extensively documented
  • Primary source materials (emails, documentation, network maps) preserved by Computer History Museum and other institutions
  • Status: Historical monument, foundational artifact

Hypercard Stacks (1987-2004)

  • Early multimedia authoring tool for Macintosh
  • Thousands of stacks created (educational software, art projects, interactive fiction)
  • Many preserved by Internet Archive’s Hypercard Stack Archive
  • Studied as precursors to the web
  • Status: Curated collection, historically significant

Early Email Archives (various)

  • Important historical emails preserved by institutions
  • Example: Jon Postel’s email about DNS root control (1980s)
  • Example: Tim Berners-Lee’s WorldWideWeb proposal (1989)
  • Status: Primary sources for internet history

Preservation Needs §

Petribytes need curatorial maintenance:

  1. Format migration: Periodically transferring to new storage media
  2. Emulation updates: Keeping emulators functional as operating systems evolve
  3. Metadata enrichment: Adding scholarly annotations, historical context
  4. Access infrastructure: Maintaining websites, databases, and discovery systems
  5. Legal protection: Ensuring copyright and ownership issues are resolved

Unlike Vivibytes (which need urgent rescue) or Umbrabytes (which are in danger of vanishing), Petribytes are institutionally secure. But they’re not invulnerable—institutions can fail, budgets can be cut, and storage media can degrade.

Ethical Considerations §

Canonization: Which artifacts become Petribytes? The selection is often biased toward:

  • Artifacts from wealthy institutions or well-documented contexts
  • Creations by famous or influential people
  • Projects with good documentation and advocacy

Meanwhile, artifacts from marginalized communities or underfunded projects often remain Umbrabytes—lost and unmourned.

Access vs. Preservation: Museums often prioritize preservation over access (artifacts locked in temperature-controlled vaults). Is this ethical? Should Petribytes be freely accessible, or is controlled access necessary for preservation?

Ownership: Who owns Petribytes? Original creators? Institutions? The public? Disputes over ownership can restrict access or lead to artifacts being removed from public view.

Triage Priority: Low (but not zero) §

Petribytes are the least urgent:

  • They’re already preserved
  • They’re institutionally supported
  • They’re documented and accessible

But they’re not safe forever:

  • Institutions can shut down
  • Budgets can be cut
  • Political shifts can lead to censorship or deaccession

Triage priority increases if:

  • The institution is unstable
  • The Petribyte is unique (no redundant copies)
  • Access is threatened (legal disputes, political pressure)

The Taxonomy in Practice: Case Study Analysis §

Let’s apply the taxonomy to a complex case: LiveJournal.

LiveJournal: A Multi-Category Artifact §

LiveJournal (founded 1999) was a blogging and social networking platform. Over its history, different parts of it occupy different taxonomic categories:

Vivibyte (1999-2017)

  • LiveJournal was alive and actively used
  • By the 2010s, it was declining but still functional
  • Users could access their posts, comments, and communities

Transition Period (2017-present)

  • LiveJournal’s Russian ownership implemented new TOS requiring compliance with Russian law
  • Many users abandoned the platform, moving to Dreamwidth or other alternatives
  • The platform is still technically alive, but English-language usage has collapsed

Archaeobyte (partial)

  • Many users exported their journals to Dreamwidth or downloaded backups
  • Internet Archive captured many public LiveJournal pages
  • Some users deleted their journals, but copies survive in archives

Umbrabyte (partial)

  • Private or friends-only journals weren’t archived by Internet Archive (respect for privacy settings)
  • Deleted journals are mostly lost (unless users saved backups)
  • Communities that were deleted by moderators often vanished without trace

Petribyte (emerging)

  • Some significant LiveJournals are being recognized as historically important:
  • Early fandom communities studied by fan studies scholars
  • Political blogs from the 2000s cited in journalism history
  • Personal journals documenting historical events (9/11, Iraq War, Arab Spring)
  • Academic papers analyze LiveJournal culture, citing preserved examples

Taxonomic Insight: LiveJournal doesn’t fit neatly into one category. Different parts of it occupy different states simultaneously. This is common with large platforms.


Taxonomy as Triage Tool §

The Archaeobyte Taxonomy isn’t just academic—it’s a practical triage tool. When deciding where to focus preservation efforts, ask:

Question 1: What mortality state is this artifact in? §

  • Vivibyte: Act now, before it dies
  • Umbrabyte: Urgent rescue, but accept that loss is likely
  • Archaeobyte: Stabilize existing preservation, prevent bit rot
  • Petribyte: Maintain and curate, but not urgent

Question 2: Is there redundancy? §

  • If an artifact is preserved by multiple institutions (e.g., in Internet Archive and Library of Congress and university archives), it’s lower priority
  • If only one fragile archive exists, priority increases

Question 3: What’s the cultural significance? §

  • High significance + Vivibyte = Critical priority
  • High significance + Umbrabyte = Urgent rescue attempt
  • High significance + Petribyte = Maintain vigilantly
  • Low significance + any state = Lower priority (harsh but necessary in triage)

Question 4: What’s the technical difficulty? §

  • Easy to preserve (static HTML) + Vivibyte = Do it now
  • Hard to preserve (complex database, proprietary format) + Vivibyte = Invest resources
  • Hard to preserve + Umbrabyte = May need to accept loss

Question 5: Are there ethical concerns? §

  • Privacy violations, consent issues, potential harm = Deprioritize or don’t preserve
  • Historical significance but ethically fraught = Preserve with restricted access

Transitions Between States §

Artifacts don’t stay in one taxonomic category forever. They transition:

Common Transitions §

Vivibyte → Archaeobyte (Successful Preservation)

  • A platform announces shutdown
  • Archivists mobilize and scrape content
  • Artifacts are preserved before servers go dark
  • Example: Vine → Internet Archive

Vivibyte → Umbrabyte (Failed Preservation)

  • A platform dies unexpectedly (no warning, or warning ignored)
  • Most content is lost
  • Only fragments survive (screenshots, partial scrapes)
  • Example: Many phpBB forums

Umbrabyte → Archaeobyte (Successful Rescue)

  • Someone finds a backup, cached copy, or forensic remnant
  • Fragments are assembled into a usable archive
  • Example: GeoCities rescue via Archive Team

Umbrabyte → Permanent Loss (Failed Rescue)

  • Fragments degrade or disappear
  • No copies exist anywhere
  • Artifact is permanently lost
  • Example: Most MySpace music from 2003-2013

Archaeobyte → Petribyte (Institutional Recognition)

  • An archived artifact gains scholarly attention
  • Institutions curate it, add metadata, make it accessible
  • It becomes part of the historical canon
  • Example: Early Hypercard stacks

Petribyte → Archaeobyte (Institutional Failure)

  • An institution shuts down or loses funding
  • Curated collection reverts to raw archive
  • Example: Rare, but possible if museums or libraries close

Undesirable Transitions (Preservation Failures) §

Archaeobyte → Umbrabyte (Bit Rot)

  • Stored files become corrupted
  • Storage media fails
  • No redundant copies exist
  • Example: Hard drives degrading in private collections

Petribyte → Archaeobyte (De-curation)

  • Budget cuts eliminate curatorial staff
  • Metadata is lost or not maintained
  • Artifacts remain preserved but lose context
  • Example: Museum collections that are “preserved” but inaccessible

Critiques and Limitations of the Taxonomy §

Critique 1: Binary Thinking §

The taxonomy implies clean categories, but reality is messy. Many artifacts are partially preserved (some Archaeobyte, some Umbrabyte). LiveJournal is alive and dead depending on which part you’re looking at.

Response: The taxonomy is a heuristic, not a rigid classification. Use it to clarify thinking, not to force artifacts into boxes.

Critique 2: Cultural Bias §

Who decides what becomes a Petribyte? The taxonomy risks reinforcing canonical hierarchies—famous people’s work gets monumentalized, marginalized communities’ work stays in limbo.

Response: This is a real problem. Archaeobytologists must actively work to diversify what gets elevated to Petribyte status. Triage should account for representational gaps.

Critique 3: Ignores Context §

An artifact’s category depends on where you are. A GeoCities site is an Archaeobyte if you know about the Archive Team torrent, but an Umbrabyte to someone who doesn’t.

Response: True. The taxonomy describes artifacts relative to preservation infrastructure. As infrastructure improves, Umbrabytes can become Archaeobytes.

Critique 4: No Category for “Never Existed” §

What about artifacts that could have been preserved but never were? The tweets that were never archived, the Snapchat videos designed to disappear?

Response: These are pre-Umbrabytes—artifacts that will become ghosts if not captured. The Vivibyte category should include them as endangered.


Expanding the Taxonomy: Proposed Sub-Categories §

Some practitioners propose additional categories:

Necrobyte (The Undead) §

Artifacts that were dead but have been resurrected:

  • Flash games made playable again via Ruffle emulator
  • GeoCities sites rebuilt and re-hosted
  • Obsolete software ported to modern systems

These are technically Archaeobytes that have been given “undead life”—functional but not native to the current era.

Cryobyte (Frozen and Waiting) §

Artifacts that are intentionally preserved in suspended animation:

  • Time capsules meant to be opened in the future
  • Long-term archives (1,000-year storage projects)
  • Artifacts preserved but deliberately not made accessible yet

These are Archaeobytes with a temporal lock—Petribytes-in-waiting.

Xenobyte (Alien and Incomprehensible) §

Artifacts so old or so alien that they’re unintelligible without extensive interpretation:

  • Code written in obsolete languages with no documentation
  • File formats with no known decoder
  • Encrypted data where the key is lost

These are Archaeobytes on the verge of becoming permanently opaque.


Practical Application: Building a Triage Matrix §

Use the taxonomy to create a triage decision matrix:

Artifact Taxonomy Redundancy Significance Difficulty Ethics Priority
Twitter archive Vivibyte High (IA + LOC) High Medium Some concerns Medium
Small Discord server Vivibyte None Low Medium Privacy issues Low
MySpace fragments Umbrabyte Low Medium Very high Consent unclear Low-Medium
Flash games Archaeobyte Medium (Flashpoint) Medium High (emulation) Mostly clear Medium
ARPANET docs Petribyte High High Low (already done) Clear Low (maintain)

This matrix helps you:

  • Compare artifacts across multiple dimensions
  • Justify triage decisions (transparency for stakeholders)
  • Identify gaps (categories with no representation)
  • Track how priorities shift over time

Conclusion: Naming the Dead §

The Archaeobyte Taxonomy gives us language for digital mortality. Before we can save artifacts, we must be able to name their states:

  • This GeoCities site is an Archaeobyte—preserved but not curated
  • That Twitter account is a Vivibyte—alive but endangered
  • Those lost MySpace songs are Umbrabytes—haunting us from beyond
  • This early ARPANET email is a Petribyte—monumentally secure

Language matters because it shapes action. When we name an artifact a Vivibyte, we acknowledge its life and its peril. When we call something an Umbrabyte, we admit it’s dying and may not be saved. When we elevate something to Petribyte, we commit institutional resources to its long-term survival.

The taxonomy isn’t just descriptive—it’s diagnostic. It tells us where to look, what to save, and how to act.

In the next chapter, we’ll explore the Archive and the Anvil—the dual practices of preservation and creation that define Archaeobytology. For now, practice identifying artifacts in the wild. Look at your own digital life. What category is your Instagram account? Your childhood blog? Your email archive?

Learn to see the world through taxonomic eyes. Because once you can name the dead, you can begin to save them.


Discussion Questions §

  1. On Categories: Choose three digital artifacts from your own life (social media profiles, old websites, photos, etc.). Classify each using the Archaeobyte Taxonomy. What does this reveal about your digital mortality?

  2. On Transitions: Describe a platform you used that transitioned from Vivibyte to Archaeobyte (or to Umbrabyte). What was that experience like? Did you try to preserve your content?

  3. On Ethics: Should we preserve Umbrabytes even when original creators might not want them resurrected? Where’s the line between historical preservation and violation of privacy?

  4. On Canonization: Why do some artifacts become Petribytes (monumentally preserved) while others remain Umbrabytes (lost and forgotten)? What biases shape this selection?

  5. On Liminal States: Can you think of an artifact that exists in multiple taxonomic states simultaneously? How does that complcomplicate preservation decisions?

  6. On Your Own Mortality: If you died tomorrow, what would happen to your digital artifacts? Would they become Archaeobytes (preserved), Umbrabytes (fragments), or simply vanish?


Exercise: Taxonomic Field Work §

Part 1: Identify and Classify

Find five digital artifacts (from your own life or the wider web) and classify each:

  1. Artifact name and URL (if applicable)
  2. Taxonomy category (Vivibyte, Archaeobyte, Umbrabyte, Petribyte)
  3. Justification (Why does it fit this category?)
  4. Transition risk (Could it move to a different category? How soon?)
  5. Preservation status (Is anyone archiving it? Where?)

Part 2: Create a Triage Matrix

Build a simple triage matrix for your five artifacts using these criteria:

  • Cultural significance (1-5 scale)
  • Endangerment level (1-5 scale)
  • Preservation difficulty (1-5 scale)
  • Ethical clarity (1-5 scale, where 5 = clearly ethical to preserve)

Part 3: Make Triage Decisions

Based on your matrix:

  • Which artifact is highest priority to preserve?
  • Which is lowest priority?
  • Are there any you would not preserve for ethical reasons?

Part 4: Reflection

Write 500 words reflecting on:

  • Did the taxonomy help you think more clearly about these artifacts?
  • Were there artifacts that didn’t fit neatly into categories?
  • How did you weigh cultural significance against endangerment level?
  • Did you discover artifacts you’d forgotten about? What was that experience like?

Further Reading §

On Digital Mortality and Preservation §

  • Chun, Wendy Hui Kyong. “The Enduring Ephemeral, or the Future Is a Memory.” Critical Inquiry 35, no. 1 (2008): 148-171.
  • Theorizes the paradox of digital “permanence” (everything is archived) and ephemerality (everything decays)

  • Kirschenbaum, Matthew. Mechanisms: New Media and the Forensic Imagination. MIT Press, 2008.

  • Foundational text on digital forensics and materiality

  • Ernst, Wolfgang. Digital Memory and the Archive. University of Minnesota Press, 2013.

  • Media archaeology perspective on digital preservation

On Platform Death §

  • Gillespie, Tarleton. “The Relevance of Algorithms.” In Media Technologies, edited by Tarleton Gillespie, Pablo Boczkowski, and Kirsten Foot, 167-194. MIT Press, 2014.
  • How platforms shape what persists and what disappears

  • Brügger, Niels. “Website History and the Website as an Object of Study.” New Media & Society 11, no. 1-2 (2009): 115-132.

  • Theorizes websites as historical objects

On Taxonomies and Classification §

  • Bowker, Geoffrey C., and Susan Leigh Star. Sorting Things Out: Classification and Its Consequences. MIT Press, 1999.
  • Classic text on how classification systems shape social reality

  • Foucault, Michel. The Order of Things. Vintage, 1994 [1966].

  • Philosophical examination of how knowledge systems are organized

On Specific Cases §

  • Brügger, Niels, and Ralph Schroeder, eds. The Web as History. UCL Press, 2017.
  • Case studies of web preservation projects

  • Ankerson, Megan Sapnar. Dot-com Design: The Rise of a Usable, Social, Commercial Web. NYU Press, 2018.

  • History of early web design, drawing on archived sites

  • Archive Team. “GeoCities: We Didn’t Start the Fire.” https://archiveteam.org/index.php?title=GeoCities

  • Primary source documenting the GeoCities rescue

End of Chapter 2

Next: Chapter 3 — The Archive and the Anvil: Dual Practices of Preservation and Creation

Part I • Theoretical Foundations & Taxonomy

Chapter 3: The Archive and the Anvil

Dual Practices of Preservation and Creation

23 min read 4,911 words

Opening: The Blacksmith and the Librarian §

Imagine two figures standing in the ruins of a murdered platform:

The Librarian surveys the wreckage with sorrow. Millions of websites, years of conversations, entire communities—all scheduled for deletion. She opens her laptop and begins downloading everything she can reach. HTML files, images, databases, user profiles. Working frantically against the shutdown clock, she fills hard drives with rescued data. When the servers go dark, she’s exhausted but determined: These artifacts will not be forgotten. I will preserve them.

The Blacksmith surveys the same wreckage with rage. Another platform murdered. Another generation of users dispossessed, their digital homes demolished by corporate landlords. He opens his laptop and begins designing. A protocol that can’t be shut down. A hosting system users can actually own. A network that survives corporate death. When the servers go dark, he’s exhausted but determined: This will not happen again. I will forge alternatives.

Both are Archaeobytologists. Both are necessary. Neither is sufficient alone.

The Archive preserves the past. The Anvil forges the future. Together, they form the dual soul of Archaeobytology—not as separate specializations, but as integrated practices that every Archaeobytologist must embody.

This chapter explores why both commitments are essential, how they complement each other, and what happens when you have one without the other.


Part I: The Archive — Practices of Preservation §

What Is the Archive? §

The Archive is not just a building full of documents. It’s a practice, a commitment, and a methodology for ensuring that the past remains accessible to the future.

In Archaeobytology, archival practice includes:

  1. Excavation: Actively rescuing artifacts before they disappear
  2. Preservation: Storing artifacts in stable, redundant, long-term formats
  3. Curation: Organizing artifacts so they’re discoverable and meaningful
  4. Interpretation: Providing context so future generations understand what they’re looking at
  5. Access: Making archives available to researchers, communities, and the public

The Archive is retrospective—it looks backward to save what’s endangered.

The Archival Impulse: Why We Save §

Why preserve murdered platforms? Why not let them die and focus only on building new ones?

Reason 1: Memory Is Identity

Communities are defined by their histories. When GeoCities died, thousands of people lost not just websites but evidence of their past selves—teenage creativity, early experiments with web design, records of online friendships from 20 years ago.

Without archives, we experience forced amnesia. Platforms control not just the present but the past. If Facebook decides to delete old posts, entire personal histories vanish. The Archive resists this erasure.

Reason 2: Cultural Continuity

Every artistic movement, every subculture, every community practice builds on what came before. Fan fiction writers today are influenced by LiveJournal fic from the 2000s. Meme culture evolves from 4chan, Tumblr, and Twitter artifacts. Web designers learn by studying archived sites from the 1990s.

If we don’t preserve digital culture, each generation starts from zero. The Archive ensures cultural continuity.

Reason 3: Historical Accountability

Archives hold powerful actors accountable. Political speeches, corporate promises, deleted tweets from public figures—these artifacts become evidence. When a politician claims they “never said that,” archived screenshots prove otherwise.

The Archive serves as collective memory against revisionism.

Reason 4: Learning from Failure

Every murdered platform teaches lessons about what went wrong:

  • Why did GeoCities users not own their domains?
  • Why couldn’t Vine users export their videos?
  • Why did Mastodon’s federation lead to fragmentation?

We can’t learn these lessons if we don’t preserve evidence. The Archive enables institutional learning.

Core Archival Practices §

1. Excavation: Rescue Before Death

The Challenge: Platforms often give little warning before shutdown—sometimes just weeks. You must act fast.

Methods:

  • Web scraping: Automated tools (wget, ArchiveBox, archive.org’s wayback-machine-downloader) download entire sites
  • API harvesting: Using platform APIs (while they still exist) to bulk-download content
  • User mobilization: Recruiting volunteers to save content manually
  • Database extraction: Obtaining database dumps from platforms (rare, requires cooperation)

Case Study: The Vine Rescue (2017)

When Vine announced shutdown, Internet Archive mobilized immediately. They:

  • Used Vine’s public API to enumerate all video IDs
  • Downloaded videos using parallel scrapers (thousands simultaneously)
  • Saved metadata (usernames, post dates, view counts, loops)
  • Captured 6.5 million videos before shutdown

Result: Vine is dead, but millions of vines survived as Archaeobytes. Researchers can study Vine culture. Creators can access their old content. Memes live on.

Lesson: Excavation requires technical skill, speed, and infrastructure (servers, bandwidth, storage).

2. Preservation: Storing for Decades

The Challenge: Digital storage degrades. Hard drives fail. File formats become obsolete. Organizations shut down. How do you preserve artifacts for 50+ years?

Strategies:

  • Redundancy: Multiple copies in multiple locations (LOCKSS principle: “Lots of Copies Keep Stuff Safe”)
  • Format migration: Periodically converting files to current standards (but risks losing fidelity)
  • Emulation: Preserving original formats + software to read them
  • Distributed storage: BitTorrent, IPFS, peer-to-peer networks where no single entity controls everything
  • Institutional partnerships: Working with libraries, universities, governments with long-term mandates

Case Study: Internet Archive’s Approach

Internet Archive maintains:

  • Primary storage: Data centers in San Francisco and Richmond, California
  • Mirror site: Complete backup in Alexandria, Egypt (Library of Alexandria partnership)
  • Glacier storage: Amazon’s long-term archival storage for redundancy
  • Partner libraries: 1,000+ libraries worldwide mirroring collections

If one data center burns down, the archive survives. If Internet Archive the organization shuts down, partner libraries can continue access.

Lesson: Preservation requires paranoia. Assume disaster. Plan for institutional failure. Build redundancy everywhere.

3. Curation: Making Sense of Data Dumps

The Challenge: Raw archives are often unusable. The Archive Team’s GeoCities torrent is 650GB of HTML files with no search function, no organization, no context.

Curation Practices:

  • Metadata creation: Adding descriptions, tags, dates, creators, context
  • Taxonomic organization: Grouping artifacts by theme, time period, community, genre
  • Search infrastructure: Building databases and search engines
  • Sampling and highlighting: Creating curated collections from massive dumps (“Best of GeoCities,” “Historically Significant Vines”)
  • Community participation: Inviting former users to add context and memories

Case Study: The 9/11 Digital Archive

After September 11, 2001, the Library of Congress and CUNY created a digital archive of:

  • Personal stories submitted by the public
  • Photos and videos from that day
  • Emails and instant messages
  • Websites created in response

This wasn’t a raw data dump. It was curated:

  • Submissions were reviewed and tagged
  • Themes were identified (first responders, survivors, international responses)
  • Oral histories were transcribed
  • Educational resources were created

Result: Not just preserved, but legible—usable by teachers, documentarians, historians, the public.

Lesson: Curation transforms data into knowledge. It’s labor-intensive but essential.

4. Interpretation: Context Is Everything

The Challenge: Future generations won’t understand artifacts without context. A GeoCities page with flashing text and tags seems bizarre now—but in 1998, it was cutting-edge design.

Interpretive Work:

  • Historical context: When was this made? What was happening politically, culturally, technologically?
  • Platform affordances: What features shaped how people communicated? (Twitter’s 140 characters, Vine’s 6 seconds)
  • Community norms: What were the unwritten rules? In-jokes? Status hierarchies?
  • Technical constraints: Why do old websites look the way they do? (Dial-up speeds, 800x600 screen resolution, limited CSS)

Case Study: Cameron’s World (GeoCities Archive)

Cameron’s World is a web art project that curates and interprets GeoCities:

  • Assembles GIFs, backgrounds, and visual elements from archived GeoCities sites
  • Presents them as a chaotic, nostalgic collage
  • Includes essays explaining GeoCities aesthetics and culture
  • Makes 1990s web design legible to people who never experienced it

This isn’t just preservation—it’s translation across time.

Lesson: Archives without interpretation become inscrutable. Future archaeologists need guides.

5. Access: Who Gets to See What?

The Challenge: Should archives be fully public? Some artifacts contain privacy violations, traumatic content, or copyrighted material.

Access Models:

Open Access (Internet Archive model)

  • Anyone can browse, search, download
  • Maximizes utility for researchers and public
  • Risk: Privacy violations, copyright disputes

Researcher Access (Library of Congress model)

  • Must apply for access, demonstrate scholarly purpose
  • Protects privacy and sensitive material
  • Risk: Limits public knowledge, creates gatekeeping

Community Access (Indigenous archives model)

  • Material is available only to the community it came from
  • Respects consent and cultural protocols
  • Risk: Limits broader historical understanding

Tiered Access (Hybrid model)

  • Public metadata (this artifact exists, here’s a description)
  • Restricted full content (apply for access)
  • Embargoes (wait X years before opening)

Case Study: Tumblr’s NSFW Purge (2018)

Tumblr banned all “adult content” in 2018, deleting millions of posts. Many were:

  • Sex education resources
  • LGBTQ+ identity expression
  • Art (nudes, erotic fiction)
  • Sex worker portfolios

Some archivists saved purged content. But should they make it public? Ethical tensions:

  • Argument for access: This is cultural heritage, representing marginalized communities
  • Argument against: Creators didn’t consent to preservation, may not want content resurrected

No easy answer. Archives must navigate these dilemmas case-by-case.

Lesson: Access is political. Every choice about who can see what shapes power and knowledge.


Part II: The Anvil — Practices of Creation §

What Is the Anvil? §

The Anvil is where we forge alternatives. It’s the practice of building tools, platforms, protocols, and institutions that embody digital sovereignty—systems designed to resist the forces that murdered previous platforms.

The Anvil is prospective—it looks forward to build what doesn’t yet exist.

The Forging Impulse: Why We Build §

Why not just preserve murdered platforms and accept that future platforms will also be murdered? Why try to build alternatives?

Reason 1: Preservation Isn’t Justice

The Archive saves artifacts, but it doesn’t change the power structures that killed them. Preserving GeoCities doesn’t give users back their domains. Archiving Vine doesn’t return ownership to creators.

The Anvil seeks systemic change—building infrastructure where users own their ground, control their data, and can’t be evicted.

Reason 2: Learning Requires Application

Studying murdered platforms teaches lessons. But those lessons are useless if we don’t apply them by building better systems. The Anvil is where theory becomes practice.

Reason 3: Alternatives Create Pressure

When people have options—federated social networks, self-hosted blogs, cooperative platforms—corporate platforms must compete. They can’t ignore user demands if users can leave.

The Anvil creates exit options that shift power dynamics.

Reason 4: Building Is Hope

Preservation is about mourning loss. Creation is about asserting possibility. The Anvil says: We don’t have to accept platform feudalism. We can forge a different future.

Core Forging Practices §

1. Tool-Making: Empowering Users

The Goal: Create software that gives people sovereignty without requiring technical expertise.

Examples:

Webrecorder (2015-present)

  • Allows anyone to archive web pages, including dynamic content (JavaScript, video embeds)
  • Runs in browser, no coding required
  • Users own their archives (WARC files they can host anywhere)
  • Sovereignty achieved: Users preserve their own history without depending on Internet Archive

Obsidian / Roam Research (2020-present)

  • Note-taking apps that store files locally in plain text (Markdown)
  • No cloud dependency (though cloud backup is optional)
  • If the company shuts down, your notes survive (unlike Evernote)
  • Sovereignty achieved: Your knowledge base isn’t hostage to a platform

Mastodon (2016-present)

  • Federated social network (anyone can run an instance)
  • ActivityPub protocol allows cross-instance communication
  • If your instance shuts down, you can migrate to another and take followers
  • Sovereignty achieved: No single corporation controls the network

Lesson: Tools should lower barriers to sovereignty. Not everyone can self-host, but tools should make it possible for those who want to.

2. Protocol Design: Building Interoperable Infrastructure

The Goal: Create open standards that allow platforms to communicate without corporate gatekeepers.

Examples:

ActivityPub (2018, W3C standard)

  • Protocol for federated social networking
  • Used by Mastodon, Pixelfed, PeerTube, and others
  • Allows users on different platforms to follow, reply, and share across networks
  • Sovereignty achieved: No single platform controls social graphs

RSS (1999, evolved through 2000s)

  • Simple protocol for syndicating content
  • Anyone can publish an RSS feed; anyone can subscribe with any reader
  • Decentralized (no company owns RSS)
  • Google Reader’s death (2013) didn’t kill RSS—new readers emerged
  • Sovereignty achieved: Publishers and readers connect directly

IPFS (InterPlanetary File System, 2015-present)

  • Peer-to-peer protocol for storing and sharing files
  • Content-addressed (files identified by hash, not location)
  • No central servers—files distributed across network
  • Sovereignty achieved: Content can’t be censored by shutting down one server

Lesson: Protocols outlive platforms. Email survived because it’s a protocol (SMTP), not a platform. Build protocols, not walled gardens.

3. Institution Building: Creating Durability

The Goal: Design organizations that can sustain preservation and sovereignty work for decades—outliving founders, surviving funding crises, resisting capture.

Examples:

Internet Archive (1996-present)

  • Non-profit with 30-year track record
  • Funded by donations, grants, and services (scanning books for libraries)
  • Governance: Board of directors, not single founder dictator
  • Mission clarity: “Universal access to all knowledge”
  • Durability factors: Diverse funding, institutional partnerships, legal advocacy (fights for fair use)

Wikimedia Foundation (2003-present)

  • Supports Wikipedia and sister projects
  • Funded by millions of small donations (avoiding capture by wealthy donors)
  • Open governance (community-elected board members)
  • Transparent financials (publishes annual reports)
  • Durability factors: Community ownership, distributed fundraising, clear mission

The Long Now Foundation (1996-present)

  • Focuses on long-term thinking (10,000-year perspective)
  • Projects include: Rosetta Project (preserving languages), 10,000-Year Clock
  • Funded by memberships, grants, and wealthy patrons who share the vision
  • Durability factors: Long time horizon built into mission, patient capital

Lesson: Institutions die from founder dependence, funding concentration, mission drift, or governance capture. Design against these failure modes from day one.

4. Designing for the Three Pillars

The Goal: Every tool, protocol, or institution should embody the Three Pillars—Declaration, Connection, Ground.

Design Questions:

Declaration (I Am)

  • Can users have persistent, self-owned identities? ([email protected], not platform/username)
  • Can they move identities between services?
  • Can they assert existence without corporate permission?

Connection (Instant Message)

  • Can users communicate directly, not through intermediaries?
  • Are relationships exportable (can you take followers/friends if you migrate)?
  • Is discovery controlled by algorithms or by users?

Ground (Digital Real Estate)

  • Do users own their data? (Can they download everything in usable formats?)
  • Do they own infrastructure? (Self-hosted, or able to migrate between hosts?)
  • Can they modify or fork the tools they use? (Open source?)

Case Study: Ghost vs. Medium

Both are blogging platforms. Compare their sovereignty:

Medium

  • Declaration: Writers get medium.com/@username (not their domain)
  • Connection: Audience belongs to Medium (can’t export email list)
  • Ground: Content is hosted on Medium servers; export is possible but clunky
  • Assessment: Low sovereignty (platform lock-in)

Ghost

  • Declaration: Writers can use custom domains (their-blog.com)
  • Connection: Audience data is exportable (email lists, subscriber data)
  • Ground: Can self-host Ghost (open source), or use Ghost(Pro) and migrate later
  • Assessment: High sovereignty (users own identity, audience, infrastructure)

Lesson: Sovereignty isn’t binary—it’s a spectrum. Ghost is more sovereign than Medium, but still less sovereign than a fully self-coded blog.

5. Resistance Architecture: Designing Against Capture

The Goal: Build systems that resist the forces that killed previous platforms—corporate acquisition, advertising pressure, venture capital extraction, government censorship.

Design Strategies:

Strategy 1: Non-Profit Structure

  • Can’t be acquired by for-profit companies
  • Mission > profit (legally required)
  • Example: Wikimedia, Internet Archive, Mozilla Foundation

Strategy 2: Cooperative Ownership

  • Users own the platform collectively
  • Decisions made democratically
  • Example: Platform cooperatives like Stocksy (photographer co-op), Resonate (musician co-op)

Strategy 3: Federated or P2P Architecture

  • No central servers to shut down
  • No single point of failure or control
  • Example: Mastodon (federated), BitTorrent (P2P), Tor (onion routing)

Strategy 4: Open Source + Copyleft

  • Code is public and forkable
  • GPL or AGPL license prevents proprietary capture
  • If maintainers sell out, community can fork
  • Example: Nextcloud (forked from ownCloud when it went proprietary)

Strategy 5: Exit Rights Built In

  • Data export is easy and complete
  • Protocols are open (can migrate to competitors)
  • No lock-in by design
  • Example: ActivityPub (can move Mastodon accounts between servers)

Case Study: WordPress’s Resistance to Capture

WordPress powers 40%+ of the web. Why hasn’t it been captured?

  • Open source: GPL-licensed, anyone can fork
  • Federated control: Core is managed by WordPress Foundation (non-profit), but thousands of independent developers contribute
  • Commercial ecosystem coexists: WordPress.com (for-profit) and WP Engine (hosting) make money, but can’t capture the open-source core
  • Portability: Easy to move WordPress sites between hosts

Result: 20+ years of survival despite corporate pressures.

Lesson: Resistance must be architected from the start. Retrofitting sovereignty into a centralized platform is nearly impossible.


Part III: Why Both Are Necessary — The Dual Soul §

The Failure of Archive-Only §

Scenario: Imagine Archaeobytology as purely preservation. We save murdered platforms but build nothing new.

What happens:

  • We accumulate vast archives of platform deaths
  • We document failure after failure
  • We become curators of a graveyard—useful for historians, but powerless to change the future
  • Each new generation experiences the same platform murders
  • We mourn endlessly but prevent nothing

This is not enough.

Archives without alternatives accept the status quo. They say: “Platforms will murder digital culture, and we’ll clean up the corpses.” That’s valuable work, but it’s defensive, reactive, and ultimately defeatist.

The Failure of Anvil-Only §

Scenario: Imagine Archaeobytology as purely creation. We build new platforms but ignore murdered ones.

What happens:

  • We repeat mistakes because we didn’t study failures
  • We reinvent the wheel, wasting effort on problems solved decades ago
  • We lose cultural continuity—each generation starts from zero
  • We abandon communities whose platforms died (no archive to return to)
  • We become techno-optimists, assuming new tools solve all problems

This is not enough.

Building without remembering is arrogant. It says: “The past doesn’t matter; we’ll build the future from scratch.” But history is full of well-meaning projects that failed because they ignored lessons of previous failures.

The Integrated Practice: Archive ⇄ Anvil §

The virtuous cycle:

  1. Study murdered platforms (Archive): What went wrong? Why did GeoCities users lose their sites?
  2. Extract lessons: Users didn’t own domains. Centralized hosting created single point of failure.
  3. Design alternatives (Anvil): Build federated hosting, encourage custom domains, create easy export tools.
  4. Document the new systems (Archive): Record how they work, why they were designed this way, what problems they solve.
  5. Iterate as systems evolve (Anvil): Improve based on user feedback and new threats.
  6. Preserve everything (Archive): Future generations can study both failures and successes.

Example: Mastodon’s Evolution

  • Archive: Studied Twitter’s centralization problems (shadowbanning, algorithmic curation, corporate control)
  • Anvil: Built Mastodon with federation (many instances, no central control)
  • Archive: Documented Mastodon’s challenges (defederation drama, moderation disputes, instance admin burnout)
  • Anvil: Improved governance (better mod tools, admin support resources)
  • Ongoing: Archive current state, forge improvements, repeat

This is the dual soul in action.

Practitioner Profiles: Embodying Both §

Not every Archaeobytologist is equally skilled at preservation and creation. But all should understand and respect both.

Profile 1: The Archivist-Who-Codes

  • Primary strength: Preservation (curation, metadata, access systems)
  • Secondary skill: Can write scrapers, build databases, maintain infrastructure
  • Example: Internet Archive staff who both curate collections and maintain the Wayback Machine

Profile 2: The Builder-Who-Preserves

  • Primary strength: Creation (software development, protocol design, system architecture)
  • Secondary skill: Understands archival needs, designs with preservation in mind
  • Example: Mastodon’s Eugen Rochko, who built a federated platform inspired by studying centralized platforms’ failures

Profile 3: The Scholar-Practitioner

  • Balances both equally: studies dead platforms, builds alternatives, publishes research
  • Example: Brewster Kahle (founded Internet Archive, advocates for digital rights, builds tools)

The key: You don’t have to be 50/50 Archive/Anvil. But you must value both and understand how they complement each other.


Part IV: Case Studies in Dual Practice §

Case Study 1: The Fediverse (Mastodon, Pixelfed, PeerTube) §

Archive Work:

  • Studied centralized social media failures (Twitter banning, Facebook surveillance, YouTube demonetization)
  • Documented what users lost when platforms changed (reach, followers, content)
  • Identified common failure modes (single corporation owns network effects)

Anvil Work:

  • Built ActivityPub protocol (open standard for federated social networking)
  • Created multiple implementations (Mastodon for microblogging, Pixelfed for photos, PeerTube for video)
  • Designed for sovereignty (users can run instances, migrate accounts, export data)

Result: Not perfect (federation has challenges—moderation complexity, discoverability issues, instance admin burnout). But represents a genuine alternative to platform capitalism.

Dual Soul Assessment: Strong Anvil (building alternatives), weaker Archive (less focus on preserving Twitter/Facebook artifacts). Could improve by integrating archived case studies into protocol design.

Case Study 2: The Internet Archive §

Archive Work:

  • Wayback Machine: 800+ billion web pages archived since 1996
  • Software collection: preserves obsolete games, applications, operating systems
  • Book digitization: scans millions of out-of-print books
  • TV and radio archives: preserves broadcast media

Anvil Work:

  • Built open-source tools (Heritrix crawler, OpenLibrary platform, Archive-It service)
  • Advocates for legal changes (fights for fair use, right to repair, library lending)
  • Supports federated archiving (encourages others to run preservation nodes)

Result: World’s most important digital preservation institution. Not just storing—actively building tools and advocating for systemic change.

Dual Soul Assessment: Strong Archive (unmatched preservation capacity), improving Anvil (tool-building and advocacy growing over time).

Case Study 3: Archive Team §

Archive Work:

  • Guerrilla archiving: scrapes dying platforms with little warning
  • Distributed effort: coordinates volunteers worldwide
  • Saves platforms institutions ignore (small forums, niche sites, “unimportant” platforms)

Anvil Work:

  • Builds scraping tools (ArchiveBot, custom scrapers for each platform)
  • Documents methodologies (how-to guides for archiving different platform types)
  • Creates preservation infrastructure (tracking systems, storage coordination)

Result: Complementary to Internet Archive—faster, more agile, less concerned with legality. Operates in gray areas institutions can’t.

Dual Soul Assessment: Strong on both Archive and Anvil. Preserves aggressively, builds tools constantly. Weakness: less focus on curation and access (creates data dumps, less interpretation).

Case Study 4: Perma.cc (Harvard Library Innovation Lab) §

Archive Work:

  • Preserves links cited in legal documents and scholarly articles
  • Prevents “link rot” in citations (URLs breaking over time)
  • Partners with law reviews, journals, and courts

Anvil Work:

  • Built simple tool: users submit URL, get permanent archive link
  • Created sustainable model: free for individuals, subscriptions for institutions
  • Designed for integration: plugins for legal citation managers

Result: Solves specific, high-value problem (preserving legal and scholarly citations). Not comprehensive like Internet Archive, but deeply integrated into academic and legal workflows.

Dual Soul Assessment: Balanced. Preserves strategically (high-value citations), builds pragmatically (easy-to-use tools), sustains institutionally (Harvard backing + subscription model).


Part V: Practical Integration — How to Embody Both §

For Individuals: Building Your Dual Practice §

If you’re primarily an archivist, add Anvil skills:

  • Learn basic coding (Python for scrapers, SQL for databases)
  • Study system design (how do resilient institutions work?)
  • Contribute to preservation tools (file bugs, write documentation, add features)

If you’re primarily a builder, add Archive skills:

  • Study platform histories (what already failed and why?)
  • Learn preservation formats (WARC, MARC, Dublin Core metadata)
  • Design with archiving in mind (build export tools, document your decisions)

For everyone:

  • Read both preservation literature and system design papers
  • Follow both archivists (e.g., @textfiles, @ArchiveTeam) and builders (e.g., @Gargron of Mastodon)
  • Contribute to projects that do both (Internet Archive, Flashpoint, Mastodon)

For Institutions: Integrating Archive and Anvil §

Museums and Libraries:

  • Don’t just preserve—build tools that others can use
  • Offer workshops on digital sovereignty (how to own your domain, self-host, export data)
  • Advocate for laws that protect both preservation and user rights

Universities:

  • Create interdisciplinary programs combining preservation, CS, law, and ethics
  • Host both archival infrastructure (servers, storage) and creation labs (makerspaces, incubators)
  • Fund research on both “how to preserve” and “how to build alternatives”

Non-Profits:

  • Balance missions: preserve and advocate for change
  • Build tools in addition to running services
  • Document everything (your own work becomes Archive material for future study)

For Communities: Collective Dual Practice §

Online communities can embody the dual soul:

  • Archive: Members back up community content (forums, Discord servers, subreddits)
  • Anvil: Migrate to more sovereign platforms when possible (self-hosted forums, federated alternatives)

Example: Reddit communities migrating to Lemmy

  • Archive: Users scrape subreddit posts before leaving
  • Anvil: Set up Lemmy instances (federated Reddit alternative)
  • Result: Community preserves history and gains sovereignty

Conclusion: The Complete Archaeobytologist §

The Archive and the Anvil are not competing priorities. They are complementary practices that reinforce each other:

  • Archives teach us what not to build (failure modes to avoid)
  • Anvils create systems worth preserving (tomorrow’s archives)
  • Archives without Anvils accept defeat
  • Anvils without Archives repeat mistakes

The complete Archaeobytologist:

  • Studies murdered platforms (Archive)
  • Designs systems that resist murder (Anvil)
  • Preserves both failures and successes (Archive)
  • Advocates for laws and norms that enable sovereignty (Anvil)
  • Teaches others to do the same (both)

You are a scholar and a smith. A custodian and a strategist. A mourner and a builder.

You do not choose between Archive and Anvil. You embody both.

In the next chapter, we’ll explore the Three Pillars in depth—the normative framework that guides both preservation and creation. These principles will show you how to evaluate whether an artifact, tool, or institution embodies digital sovereignty.

For now, consider: What are you preserving? What are you building? And how do those practices reinforce each other?

The dual soul awaits.


Discussion Questions §

  1. On Personal Practice: Which role feels more natural to you—Archivist or Blacksmith? What would it take to develop skills in the other domain?

  2. On Institutional Models: Compare Internet Archive (non-profit preservation) and Mastodon (federated protocol). Which model is more sustainable long-term? Why?

  3. On Priorities: If you had to choose between (A) perfectly preserving one murdered platform or (B) building a tool that prevents future platform murders, which would you choose? Why?

  4. On Integration: Can you think of a project that successfully integrates Archive and Anvil? What does it do well? What could be improved?

  5. On Failure Modes: What happens when preservation work is done without creation? When creation happens without preservation? Find real-world examples.

  6. On Your Own Life: Audit your digital life. What are you preserving (backups, exports, archives)? What are you building (websites, tools, contributions to open platforms)?


Exercise: Design a Dual-Practice Project §

Scenario: Choose a currently-living platform you use (Twitter/X, Instagram, TikTok, Reddit, Discord, etc.). Design a project that embodies both Archive and Anvil:

Part 1: Archive Component (500 words)

  • What would you preserve from this platform?
  • How would you collect it (scraping, API, user exports)?
  • What metadata would you capture?
  • How would you organize it (taxonomy, search, curation)?
  • What ethical issues arise (privacy, consent, copyright)?

Part 2: Anvil Component (500 words)

  • What lessons does this platform teach about failure modes?
  • What alternative would you build to avoid those failures?
  • How would it embody the Three Pillars (Declaration, Connection, Ground)?
  • What technologies would you use (federated, P2P, blockchain, self-hosted)?
  • How would you ensure long-term sustainability?

Part 3: Integration (300 words)

  • How do Archive and Anvil components reinforce each other?
  • Would you preserve the old platform’s content in the new system?
  • How would you tell the story of “why we built this alternative”?
  • What would you document for future Archaeobytologists studying your work?

Part 4: Reflection (200 words)

  • Which was harder to design—Archive or Anvil?
  • Did designing one inform the other?
  • Would you actually want to undertake this project? Why or why not?

Further Reading §

On Archives and Memory §

  • Derrida, Jacques. Archive Fever: A Freudian Impression. University of Chicago Press, 1996.
  • Philosophical meditation on archives, memory, and destruction

  • Manoff, Marlene. “Theories of the Archive from Across the Disciplines.” Portal: Libraries and the Academy 4, no. 1 (2004): 9-25.

  • Survey of how different fields theorize archives

  • Cook, Terry. “What is Past is Prologue: A History of Archival Ideas Since 1898, and the Future Paradigm Shift.” Archivaria 43 (1997): 17-63.

  • Evolution of archival theory and practice

On Building Alternatives §

  • Benkler, Yochai. The Wealth of Networks. Yale University Press, 2006.
  • Theory of peer production and commons-based alternatives

  • Doctorow, Cory. The Internet Con: How to Seize the Means of Computation. Verso, 2023.

  • Advocacy for interoperability and user sovereignty

  • Schneider, Nathan. “An Internet of Ownership: Democratic Design for the Online Economy.” The Sociological Review 68, no. 2 (2020): 320-340.

  • Platform cooperatives and ownership models

On Dual Practice §

  • Kahle, Brewster. “Preserving the Internet.” Scientific American 276, no. 3 (1997): 82-83.
  • Internet Archive founder on preservation imperatives

  • Star, Susan Leigh, and Karen Ruhleder. “Steps Toward an Ecology of Infrastructure.” Information Systems Research 7, no. 1 (1996): 111-134.

  • How infrastructure shapes what can be preserved and built

  • Sennett, Richard. The Craftsman. Yale University Press, 2008.

  • Philosophy of making and building with care

Primary Sources §

  • Internet Archive. “About the Internet Archive.” https://archive.org/about/
  • Archive Team. “Who We Are.” https://archiveteam.org/
  • ActivityPub. W3C Recommendation. https://www.w3.org/TR/activitypub/
  • Perma.cc. “About Perma.cc.” https://perma.cc/about

End of Chapter 3

Next: Chapter 4 — The Three Pillars of Digital Sovereignty: Declaration, Connection, Ground

Part I • Theoretical Foundations & Taxonomy

Chapter 4: The Three Pillars of Digital Sovereignty

Declaration, Connection, Ground

24 min read 5,143 words

Opening: The Ghost in the Machine §

In 2007, a woman named Sara lost her husband to cancer. For months afterward, she found comfort in reading through their old emails—thousands of messages spanning 15 years of marriage. Love letters, vacation plans, inside jokes, mundane logistics that now felt precious. The emails were stored in her AOL account, which she’d had since 1996.

In 2013, AOL announced it would delete inactive email accounts. Sara’s husband’s account had been inactive for six years. She frantically tried to log in to save his emails, but she’d never known his password. AOL’s customer service said they couldn’t help—policy was policy. On the deletion date, every email her husband had ever sent vanished.

Sara’s husband had no Declaration—his identity was AOL’s property, revocable at their discretion. Their conversations had no Connection—all communication was mediated and stored by a corporation. They had no Ground—the emails lived on AOL’s servers, subject to AOL’s rules.

When the servers deleted his account, it was as if he’d never existed.

This is what happens when we build our digital lives on platforms we don’t own. We become tenants in digital space, vulnerable to eviction at any moment. Our identities, relationships, and memories exist only as long as corporations permit them to.

The Three Pillars offer an alternative vision: a model for digital existence where you own your identity, control your connections, and possess your ground. Not as a tenant, but as a sovereign.

This chapter explores each Pillar in depth—what it means, why it matters, and how to achieve it.


The Three Pillars: Origins and Philosophy §

Philosophical Roots §

The Three Pillars draw on multiple intellectual traditions:

1. Property Rights (Locke, Rousseau)

  • John Locke: You own the product of your labor; your body and mind are your property
  • Applied to digital: Content you create, relationships you build, data you generate—these should be yours

2. Autonomy (Kant)

  • Immanuel Kant: Rational beings deserve self-governance; autonomy is prerequisite for dignity
  • Applied to digital: You should control your digital existence without corporate intermediation

3. Sovereignty (Political Philosophy)

  • Westphalian sovereignty: States have supreme authority within their borders
  • Applied to digital: Individuals should have supreme authority within their digital domains

4. The Commons (Ostrom)

  • Elinor Ostrom: Communities can self-govern shared resources without privatization or state control
  • Applied to digital: Digital infrastructure can be collectively owned without corporate capture

Contemporary Influences §

Cory Doctorow: “Adversarial Interoperability”

  • Users should be able to modify, extend, and migrate away from platforms
  • Platforms shouldn’t be able to lock users in with technical or legal barriers

Lawrence Lessig: “Code Is Law”

  • Digital architecture shapes behavior and power
  • We must build infrastructure that embodies our values

Bruce Schneier: “Feudal Security”

  • Modern platforms create “feudal” relationships—we depend on corporate lords for protection
  • We should build systems where security doesn’t require surrendering autonomy

Shoshana Zuboff: “Surveillance Capitalism”

  • Platforms extract behavioral data as raw material for profit
  • Sovereignty requires breaking free from extraction economics

The Three Pillars as Synthesis §

The Three Pillars synthesize these ideas into a practical framework:

  1. Declaration (I Am): Self-owned identity and voice
  2. Connection (Instant Message): Direct, unmediated relationships
  3. Ground (Digital Real Estate): Owned infrastructure and data

Together, they define digital sovereignty—the ability to exist, communicate, and build in digital space without corporate gatekeeping.


Pillar 1: Declaration (I Am) §

Core Principle §

You should be able to declare your identity and existence without permission from any platform or intermediary.

Your name, your voice, your presence—these should be self-originating, not granted by Facebook, Twitter, or Google.

What Declaration Means in Practice §

Identity Ownership

  • Your username/identity is not tied to a platform: [email protected], not [email protected]
  • You control authentication: you decide who can verify you are who you claim to be
  • Persistence: your identity survives platform shutdowns

Voice

  • You can publish thoughts without platform censorship (though not freedom from legal or social consequences)
  • You control your archive: everything you’ve ever said remains accessible to you
  • No algorithmic suppression: platforms can’t shadowban or throttle your reach

Presence

  • You can be found without relying on platform search or directories
  • Your digital “home” (website, profile, portfolio) exists independently
  • You can choose to be ephemeral or permanent on your own terms

Historical Context: How We Lost Declaration §

Era 1: Early Internet (1990s)

  • People owned domains (yourname.com)
  • Email was federated (anyone could run a mail server)
  • Personal homepages were the norm
  • Declaration was default

Era 2: Platform Consolidation (2000s-2010s)

  • Social media centralized identity (Facebook profiles, Twitter handles)
  • Email became dominated by Gmail, Yahoo, Outlook
  • “Real name” policies forced legal names, erasing pseudonymous freedom
  • Declaration was lost

Era 3: Attempted Reclamation (2010s-present)

  • IndieWeb movement: reclaim your domain, own your content
  • Federated platforms: Mastodon, Matrix, ActivityPub
  • Decentralized identity: blockchain-based names, DIDs (Decentralized Identifiers)
  • Declaration is contested

Case Study: The Real Name Policy Wars §

Facebook’s Real Name Policy (2014)

  • Requirement: use legal name on profile
  • Enforcement: accounts suspended if names deemed “fake”
  • Impact: Disproportionately harmed:
  • LGBTQ+ people using chosen names
  • Abuse survivors hiding from stalkers
  • Activists in authoritarian countries
  • Native Americans with non-Western naming conventions
  • Drag performers and artists with stage names

Community Response

  • Protests, petitions, media campaigns
  • Alternative platforms emerged (Ello, Mastodon)
  • Facebook eventually softened policy but never fully reversed

Sovereignty Analysis

  • Facebook claimed authority to define “real” identity
  • Users who didn’t comply lost Declaration—couldn’t exist on platform under chosen name
  • Alternative: If users owned domains, they’d declare identity themselves (no platform veto)

Case Study: Twitter Handle Squatting and Seizure §

The Problem

  • Desirable Twitter handles (@God, @Music, @Tech) often registered early by random users
  • Companies and celebrities wanted those handles
  • Twitter could seize handles and reassign them (with or without compensation)

Examples

  • @Music: taken from a user and given to a music industry account
  • Short handles: forcibly renamed to free up namespace for corporate use
  • Parody accounts: suspended without appeal when targets complained

Sovereignty Analysis

  • Twitter usernames are leased, not owned
  • Platform can revoke at any time
  • True Declaration would mean: @[email protected] (federated identity, like email)
  • No platform could seize your identity if you own the domain

Achieving Declaration: Practical Steps §

Step 1: Own a Domain

  • Register a domain name ($10-15/year)
  • This becomes your permanent digital address
  • Even if hosting changes, the domain remains yours

Step 2: Use Domain-Based Identity

Step 3: Self-Host or Use Portable Hosting

  • Self-host if you have technical skill (full control)
  • Or use hosting you can migrate from (WordPress, Ghost, static site hosts)
  • Avoid platforms where your identity is tied to their domain (Medium.com/@you, Facebook.com/you)

Step 4: Archive Everything You Publish

  • Keep local copies of all content
  • Export data regularly from any platforms you use
  • Your archive proves you said what you said (even if platforms delete it)

Spectrum of Sovereignty

Platform Identity Portability Control Declaration Score
Facebook facebook.com/you None Platform ★☆☆☆☆
Twitter @you None Platform ★☆☆☆☆
Medium medium.com/@you Export possible Platform ★★☆☆☆
Ghost you.ghost.io or custom domain Full export Hybrid ★★★☆☆
Mastodon (hosted) @[email protected] Account migration Instance admin ★★★☆☆
Mastodon (own instance) @[email protected] Full You ★★★★☆
Self-hosted site yourdomain.com Full You ★★★★★

Critiques and Limitations §

Critique 1: “Not everyone can afford domains”

  • Domains cost $10-15/year—not free, but not prohibitive for many
  • Possible solutions: Subsidized domains for low-income users, community domain cooperatives

Critique 2: “Most people don’t want to manage infrastructure”

  • True—self-hosting requires technical skill and time
  • Compromise: Use platforms that support custom domains (Ghost, WordPress)
  • Still achieves Declaration (own your identity) without full self-hosting

Critique 3: “Domains can be seized too” (government, ICANN, registrars)

  • Valid concern—DNS is centralized and vulnerable
  • Alternative solutions: Blockchain-based names (ENS, Namecoin), though these have their own problems (cost, complexity)
  • No system is perfectly sovereign, but domains are more sovereign than platform usernames

Critique 4: “Pseudonymity is harder with domains”

  • Domains require registration (name, address, though WHOIS privacy helps)
  • Platform pseudonyms (Twitter handles) are easier for anonymity
  • Trade-off: sovereignty vs. anonymity
  • Possible solution: Domains registered through privacy-preserving services or cooperatives

Pillar 2: Connection (Instant Message) §

Core Principle §

You should be able to communicate directly with others without a platform mediating, monitoring, or monetizing your relationships.

Your connections—friendships, communities, audiences—should be portable and platform-independent, not locked inside corporate silos.

What Connection Means in Practice §

Direct Communication

  • Messages go peer-to-peer or through neutral infrastructure (not corporate servers logging everything)
  • No algorithmic filtering: if you send a message, recipient sees it (unless they block you)
  • No surveillance: platforms don’t read your messages for advertising or AI training

Portable Relationships

  • Your “social graph” (who you follow, who follows you) is exportable
  • If you leave a platform, you can take your connections with you
  • Relationships aren’t held hostage by network effects

Intentional Discovery

  • You choose who to connect with (not algorithmic recommendations)
  • Communities form organically, not through platform-engineered “engagement”
  • No shadow manipulation (algorithmic amplification/suppression invisible to users)

Historical Context: How We Lost Connection §

Era 1: Email and Forums (1990s-2000s)

  • Email was federated: Gmail users could email Outlook users
  • Forums were independent: each community ran its own servers
  • IRC, XMPP: open protocols for chat
  • Connection was open and portable

Era 2: Platform Silos (2000s-2010s)

  • Social media created walled gardens: Facebook users couldn’t message Twitter users
  • Network effects locked users in: everyone’s on Facebook, so you have to be too
  • Algorithmic feeds: platforms decided what you see (not chronological)
  • Connection was enclosed and mediated

Era 3: Attempted Reopening (2010s-present)

  • Federated social media: ActivityPub (Mastodon, Pixelfed, Lemmy)
  • End-to-end encryption: Signal, Matrix, secure messaging
  • Interoperability advocacy: EU’s Digital Markets Act requires platform interoperability
  • Connection is being contested

Case Study: Facebook’s Closed Graph §

The Problem

  • Facebook has 3 billion users—largest social graph in history
  • You can’t export your social graph (list of friends/followers is platform-locked)
  • Can’t communicate with friends on other platforms (Instagram, Twitter, Mastodon)
  • If you leave Facebook, you lose access to your network

Example: The 2021 Exodus

  • Concerns over privacy, misinformation, mental health led some users to quit Facebook
  • But: leaving meant losing contact with family, community groups, event organizing
  • Many felt trapped: “I hate Facebook, but I can’t leave because everyone’s there”

Sovereignty Analysis

  • Facebook owns your relationships (not you)
  • Network effects create economic lock-in: cost of leaving is too high
  • True Connection would mean: export your friends list, communicate with them on any platform

What Sovereignty Would Look Like

  • You export friend list with contact info: emails, domain-based identities
  • You follow @[email protected] from any ActivityPub client
  • If you switch platforms, you import connections (like changing email clients)

Case Study: Twitter’s Algorithmic Feed §

The Problem

  • Twitter replaced chronological timeline with algorithmic feed (2016)
  • Algorithm decides what you see (optimizing for “engagement”)
  • Result: rage-bait and controversy amplified, nuanced discussions buried

User Impact

  • You follow someone, but don’t see their tweets (algorithm filtered them out)
  • They don’t even know you didn’t see it (shadow suppression)
  • Your voice is throttled invisibly (tweets shown to fewer followers)

Sovereignty Analysis

  • Platform mediates Connection—you don’t directly communicate with followers
  • Algorithm decides who sees what (no transparency, no user control)
  • True Connection would mean: chronological feed, or user-chosen filters (not platform-imposed)

Case Study: WhatsApp’s End-to-End Encryption (Partial Sovereignty) §

What WhatsApp Did Right

  • End-to-end encryption: messages can’t be read by WhatsApp servers
  • Signal Protocol: open-source, audited, gold standard for security
  • Result: private, direct communication (no platform surveillance)

What WhatsApp Still Controls

  • Metadata: who messages whom, when, how often (not encrypted)
  • Account tied to phone number (not portable identity)
  • Closed platform: can’t message Signal or Matrix users
  • Facebook acquisition: company owns platform, could change policies

Sovereignty Analysis

  • Strong on privacy (encryption)
  • Weak on portability (can’t take contacts to other platforms)
  • Partial Connection: direct communication, but within closed ecosystem

Better Model: Matrix

  • Federated protocol (like email): anyone can run a server
  • End-to-end encryption by default
  • Interoperable: message users on any Matrix server from any Matrix client
  • Account migration: can switch servers and keep contacts

Achieving Connection: Practical Steps §

Step 1: Use Federated Platforms

  • Mastodon (social media), Matrix (chat), email (already federated)
  • Can communicate across servers/instances
  • Not locked into one provider

Step 2: Export Your Social Graph Regularly

  • Download follower lists, friend lists, contact exports from platforms
  • Store locally with contact info (emails, domains, federated handles)
  • If platform dies or you leave, you can reconnect elsewhere

Step 3: Use Open Protocols

  • Email, RSS, ActivityPub, Matrix—protocols anyone can implement
  • Avoid proprietary platforms that don’t interoperate (Instagram, Snapchat)

Step 4: Support Interoperability Legislation

  • EU’s Digital Markets Act requires large platforms to interoperate
  • In US, advocate for similar laws
  • Interoperability makes it possible to leave platforms without losing connections

Spectrum of Sovereignty

Platform Communication Graph Portability Interoperability Connection Score
Facebook Messenger Mediated, surveilled None None ★☆☆☆☆
WhatsApp E2E encrypted Phone number only None ★★☆☆☆
Twitter DMs Mediated, surveilled Export limited None ★☆☆☆☆
Signal E2E encrypted Phone number Signal-only ★★★☆☆
Email Direct or federated Address book exportable Full (SMTP) ★★★★☆
Matrix E2E encrypted, federated Exportable Full (Matrix protocol) ★★★★★
Mastodon Federated, public Account migration Full (ActivityPub) ★★★★☆

Critiques and Limitations §

Critique 1: “Network effects make leaving impossible”

  • True—if everyone’s on Facebook, switching to Mastodon means losing reach
  • Solution requires critical mass: enough people must switch together
  • Interoperability laws help: if Facebook had to let you message from Mastodon, leaving wouldn’t mean disconnection

Critique 2: “Federated platforms have moderation problems”

  • Valid—federation complicates moderation (who decides what’s acceptable?)
  • Instance admins must defederate from toxic servers, creating fragmentation
  • Trade-off: sovereignty vs. ease of moderation
  • Ongoing challenge for federated systems

Critique 3: “Privacy and portability can conflict”

  • Making social graphs exportable could enable harassment (exporting someone else’s follower list to target them)
  • Solution: Export your own connections only, not others’ data about you
  • Balance: your sovereignty shouldn’t violate others’ privacy

Critique 4: “Most people prioritize convenience over sovereignty”

  • Accurate—Facebook Messenger is easier than running a Matrix server
  • Doesn’t mean we should abandon sovereignty, but signals need for user-friendly sovereign tools
  • Success case: Signal (E2E encryption as simple as WhatsApp)

Pillar 3: Ground (Digital Real Estate) §

Core Principle §

You should own the infrastructure your digital life is built on—not rent it from a landlord who can evict you.

Your data, your files, your websites, your history—these should exist on ground you control, portable and independent from any single platform’s survival.

What Ground Means in Practice §

Data Ownership

  • You can download everything: posts, photos, messages, metadata, in usable formats (not locked PDFs)
  • Data is yours legally (not “licensed” to platform)
  • You can delete permanently (right to erasure, not just “soft delete”)

Infrastructure Control

  • Self-hosted (you run the servers) or portable hosting (can migrate)
  • No platform lock-in: if provider shuts down, you move elsewhere
  • Can fork/modify tools (open source preferred)

Persistence

  • Your domain survives company shutdowns
  • URLs remain stable (no link rot from platform restructuring)
  • Content persists as long as you pay hosting/domain costs (not at platform’s whim)

Historical Context: How We Lost Ground §

Era 1: Personal Ownership (1990s)

  • Personal websites on ISP-provided space
  • Owned your files (stored locally, uploaded to server)
  • Ground was yours (within limits—still renting server space)

Era 2: Platform Enclosure (2000s)

  • MySpace, Facebook, GeoCities: free hosting in exchange for ads
  • Content lived on platform servers (not your local machine)
  • Terms of Service granted platforms broad rights to your content
  • Ground was enclosed

Era 3: The Cloud (2010s)

  • Everything in cloud: photos (Google Photos), documents (Google Docs), files (Dropbox)
  • Convenience: access from any device
  • Cost: data lives on company servers, subject to their policies
  • Ground was fully abstracted (you don’t know where your data physically is)

Era 4: Reclamation Movements (2010s-present)

  • Self-hosting: Nextcloud, Syncthing, Home servers
  • Decentralized storage: IPFS, BitTorrent, blockchain storage
  • Right-to-download laws: GDPR requires data portability
  • Ground is being contested

Case Study: GeoCities as Loss of Ground §

What Happened

  • GeoCities gave users free webspace: geocities.com/neighborhood/username
  • Users built websites, thinking they owned them
  • 2009: Yahoo shut down GeoCities with minimal warning
  • 30 million sites vanished

Why It Happened

  • Users didn’t own domains—addresses were hierarchical under geocities.com
  • Hosting was free but at Yahoo’s discretion
  • No contractual right to persistence
  • No easy way to migrate (no domain portability)

Sovereignty Analysis

  • Users had no Ground—they were digital tenant farmers
  • When landlord (Yahoo) demolished the land, they lost everything
  • True Ground would mean: own domain, portable hosting, local backups

What Could Have Prevented This

  • If users had registered domains (yourname.com) pointing to GeoCities hosting
  • When Yahoo shut down, users could’ve moved to new hosting (same domain)
  • Content would’ve survived platform death

Case Study: Google Photos’ Unlimited Storage Reversal §

The Bait

  • 2015: Google Photos launches with “free unlimited storage” (at reduced quality)
  • Millions of users upload entire photo libraries
  • Primary copies deleted from local devices (trusting cloud)

The Switch

  • 2021: Google announces unlimited storage ending
  • Users must pay or delete photos
  • Photos hostage: can’t easily migrate to other platforms (bulk download is cumbersome)

Sovereignty Analysis

  • Users lost Ground by deleting local copies
  • Google owns physical storage and can change terms
  • True Ground would mean: keep primary copies locally, use cloud only as backup
  • Or: Use distributed storage (no single company controls it)

Case Study: The Notion Migration Crisis §

Background

  • Notion: popular note-taking/project-management app
  • Users store everything in Notion: notes, projects, knowledge bases
  • Cloud-based: data lives on Notion’s servers

The Fear

  • If Notion shuts down, goes bankrupt, or gets acquired and killed—all data lost?
  • Export exists (Markdown/HTML) but imperfect (complex databases don’t export cleanly)

User Response

  • Anxiety about lock-in
  • Some users migrate to Obsidian (local Markdown files)
  • Others accept risk for convenience

Sovereignty Analysis

  • Notion users have weak Ground (data exportable but dependent on company survival)
  • Obsidian users have strong Ground (local files, company could die and files remain)
  • Trade-off: features/collaboration vs. sovereignty

Case Study: The IndieWeb Movement (Ground Reclamation) §

Principles

  1. Own your domain: yourname.com is your identity
  2. Own your content: original posts on your site (syndicate to platforms if you want reach)
  3. Own your data: keep local backups, use open formats

Practices

  • POSSE (Post On your Site, Syndicate Elsewhere): Write blog post, auto-post to Twitter/Mastodon
  • Webmentions: decentralized “comments” system (sites can reply to each other without centralized platform)
  • Micropub: protocol for publishing to your own site from any client

Example: A Day in the IndieWeb Life

  1. Write blog post on your-domain.com
  2. Auto-syndicate to Twitter, Mastodon, Reddit
  3. Replies on those platforms appear as comments on your blog (via webmention)
  4. If platforms die, your original post survives (on your domain)
  5. If you switch hosting, same domain works (portability)

Sovereignty Assessment

  • Full Ground: own domain, own data, portable hosting
  • Strong Declaration: yourname.com is persistent identity
  • Moderate Connection: can syndicate to platforms for reach, but primary home is yours

Achieving Ground: Practical Steps §

Level 1: Renters with Good Backups

  • Use platforms (Facebook, Twitter, Notion) but export data regularly
  • Keep local copies of everything important
  • If platform dies, you have your data

Level 2: Portable Tenants

  • Use platforms that support data portability and custom domains
  • WordPress, Ghost, Netlify, Vercel: can migrate to other hosting
  • Own domain, so URLs persist across migrations

Level 3: Self-Hosted Sovereigns

  • Run your own servers (VPS, home server)
  • Use open-source software (WordPress, Nextcloud, Mastodon)
  • Full control over data and infrastructure

Level 4: Distributed Ground

  • Use peer-to-peer or blockchain storage (IPFS, Filecoin, Arweave)
  • Content persists even if you disappear (no single point of failure)
  • Censorship-resistant (no entity can delete content)

Spectrum of Sovereignty

Platform Data Ownership Export Quality Domain Control Ground Score
Facebook Platform license Limited HTML None ★☆☆☆☆
Twitter Platform license JSON export None ★★☆☆☆
Medium Retain rights Markdown export None ★★☆☆☆
Ghost (hosted) You own Full export Custom domain ★★★★☆
WordPress (self-hosted) You own Full (database) Your domain ★★★★★
Static site (Netlify/Vercel) You own (in Git) Full Your domain ★★★★★
IPFS-hosted site Distributed Full Your domain + content hash ★★★★★

Critiques and Limitations §

Critique 1: “Self-hosting is too technical for most people”

  • Accurate—requires server management, security updates, backups
  • Counter: Tools are getting easier (Yunohost, Sandstorm, managed hosting)
  • Compromise: Use portable hosting (Ghost, WordPress) with custom domain (achieves most sovereignty)

Critique 2: “Distributed storage is expensive/slow”

  • Valid—IPFS/blockchain storage costs money, slower than centralized cloud
  • Counter: Costs are dropping, speeds improving
  • Use case: For critical archival content (doesn’t need daily access), distributed storage is viable

Critique 3: “My domain can still be seized”

  • True—ICANN, governments, registrars can revoke domains
  • Mitigations: Use privacy-friendly registrars, blockchain domains (ENS), Tor onion services
  • No perfect solution, but domains are more sovereign than platform URLs

Critique 4: “What about backup redundancy?”

  • Self-hosters must maintain their own backups (not automatic like Google Photos)
  • Risk: Home server fails, data lost
  • Solution: Hybrid approach (self-host primary, backup to cloud, or use distributed backup)

The Three Pillars in Practice: Sovereignty Audit §

How to Audit Your Own Sovereignty §

For each part of your digital life, ask:

Declaration:

  • Do I own my identity? (custom domain vs. platform username)
  • Can I prove I said what I said? (archive)
  • Can my identity be revoked? (platform TOS)

Connection:

  • Can I export my social graph? (follower list, friend list)
  • Can I message people on other platforms? (interoperability)
  • Are my conversations surveilled? (E2E encryption)

Ground:

  • Do I own my data? (legally and practically)
  • Can I export everything? (download in usable format)
  • Can I migrate without losing URLs? (custom domain)

Example Audit: Personal Blog §

Aspect Platform Blog (Medium) Sovereign Blog (Self-hosted WordPress)
Declaration medium.com/@username (not yours) yourdomain.com (yours)
Identity portability None Full (domain stays)
Archive control Platform can delete You control
Connection Medium network only RSS, email newsletter, federated
Reader relationships Platform-mediated Direct (email subscribers)
Discovery Medium algorithm SEO, RSS, direct links
Ground Data licensed to Medium You own
Export Markdown (good) Full database (perfect)
Persistence Medium’s discretion As long as you pay hosting

Sovereignty Score:

  • Medium: ★★☆☆☆ (some portability, but limited sovereignty)
  • Self-hosted: ★★★★★ (full sovereignty)

Example Audit: Social Media Presence §

Aspect Facebook Mastodon (own instance)
Declaration facebook.com/username @[email protected]
Identity ownership Facebook’s Yours (via domain)
Account seizure risk High (TOS violations) Low (you control server)
Connection Facebook only ActivityPub (any compatible platform)
Friend portability None Account migration
E2E encryption Messenger has it Depends on instance config
Ground Data on Facebook servers Data on your server
Export quality Limited JSON Full database
Control Facebook’s rules Your rules (your instance)

Sovereignty Score:

  • Facebook: ★☆☆☆☆ (minimal sovereignty)
  • Mastodon (own instance): ★★★★★ (high sovereignty, though federated)

Building Systems That Embody the Three Pillars §

Design Checklist for Sovereign Systems §

When building tools, platforms, or institutions, ask:

Declaration:

  • [ ] Do users control their identities? (domain-based or self-generated, not platform-assigned)
  • [ ] Can identities migrate between providers?
  • [ ] Are identities persistent (survive platform changes)?

Connection:

  • [ ] Can users communicate without platform surveillance?
  • [ ] Is the social graph exportable?
  • [ ] Does the system interoperate with other platforms? (open protocols)

Ground:

  • [ ] Do users own their data legally?
  • [ ] Can users export everything in usable formats?
  • [ ] Can users self-host, or easily migrate between hosts?

Case Study: Matrix Protocol (High Sovereignty) §

Declaration:

  • Identity: @user:homeserver.com (federated, like email)
  • You choose homeserver (or run your own)
  • Identity migrates if you change servers

Connection:

  • End-to-end encryption by default
  • Federated: message users on any Matrix homeserver
  • Social graph: friends list portable

Ground:

  • Open protocol (anyone can implement)
  • Self-hosting supported
  • Full data export

Sovereignty Score: ★★★★★ (all three pillars strong)

Case Study: ENS (Ethereum Name Service) (Partial Sovereignty) §

Declaration:

  • Own yourname.eth forever (NFT ownership)
  • Censorship-resistant (no ICANN or government can seize)
  • Can point to websites, wallets, social profiles

Connection:

  • Doesn’t directly provide communication (just naming)
  • But: can be used as identity for federated systems

Ground:

  • Own the name (on blockchain)
  • But: Expensive (initial registration + renewal gas fees)
  • And: Requires crypto wallet (technical barrier)

Sovereignty Score: ★★★☆☆ (strong Declaration, neutral Connection, weak Ground due to cost/complexity)


Conclusion: The Architecture of Freedom §

The Three Pillars aren’t just philosophical ideals—they’re design principles for building a different kind of digital future.

Every platform murder, every account suspension, every data breach is a failure of sovereignty. These crises happen because we’ve built digital infrastructure on feudal principles: users as tenants, platforms as landlords.

The Three Pillars offer an alternative:

  • Declaration: You own your name
  • Connection: You control your relationships
  • Ground: You possess your infrastructure

Together, they constitute digital freedom—not as abstract right, but as practical architecture.

In the next chapter, we’ll explore Triage Methodology—how to decide what to save when everything is endangered. The Three Pillars will guide these decisions: artifacts and systems that embody sovereignty deserve prioritization.

For now, audit your own digital life. Where do you have Declaration? Connection? Ground? And where are you vulnerable—a tenant on borrowed land, subject to eviction at any moment?

The architecture of freedom begins with seeing the chains. And then, systematically, building your way out.


Discussion Questions §

  1. Personal Audit: Conduct a Three Pillars audit of your primary digital platforms (social media, email, cloud storage, blog). Where are you sovereign? Where are you vulnerable?

  2. Trade-offs: Would you accept less convenience for more sovereignty? What’s the breaking point? (e.g., self-host email vs. use Gmail)

  3. Collective Action: Can individual sovereignty exist without collective action? If everyone stays on Facebook, does your Mastodon account matter?

  4. Privilege: Is digital sovereignty a luxury for technical elites? How do we make it accessible to everyone?

  5. Necessity: Are the Three Pillars truly necessary? Can you be “free enough” using corporate platforms with good export tools?

  6. Future Scenario: Imagine 2035. What does a maximally sovereign digital life look like? What compromises remain?


Exercise: Design a Sovereign Alternative §

Task: Choose a platform you currently use (Twitter, Instagram, Notion, Discord, etc.). Design a sovereign alternative that embodies all Three Pillars.

Part 1: Critique Current Platform (500 words)

  • How does the current platform fail each Pillar?
  • What specific sovereignty violations matter most?
  • What would users lose if the platform died tomorrow?

Part 2: Design Alternative (1000 words)

  • Declaration: How do users own their identities?
  • Connection: How do they communicate? Is it interoperable?
  • Ground: How is data stored? Who owns infrastructure?
  • What technologies enable this? (federation, P2P, blockchain, self-hosting, etc.)

Part 3: Adoption Strategy (500 words)

  • How do you get users to switch? (Network effects are powerful)
  • What’s the minimum viable product?
  • How do you sustain the system long-term? (funding, governance)

Part 4: Reflect (300 words)

  • What compromises did you make? (Perfect sovereignty is often impractical)
  • What did you learn about the tensions between convenience and sovereignty?

Further Reading §

On Digital Sovereignty §

  • Schneier, Bruce. Data and Goliath: The Hidden Battles to Collect Your Data and Control Your World. W.W. Norton, 2015.
  • Véliz, Carissa. Privacy Is Power: Why and How You Should Take Back Control of Your Data. Melville House, 2020.
  • Doctorow, Cory. The Internet Con: How to Seize the Means of Computation. Verso, 2023.

On Infrastructure and Architecture §

  • Lessig, Lawrence. Code: Version 2.0. Basic Books, 2006.
  • Star, Susan Leigh. “The Ethnography of Infrastructure.” American Behavioral Scientist 43, no. 3 (1999): 377-391.
  • Winner, Langdon. “Do Artifacts Have Politics?” Daedalus 109, no. 1 (1980): 121-136.

On Property and Ownership §

  • Locke, John. Second Treatise of Government [1689].
  • Ostrom, Elinor. Governing the Commons. Cambridge University Press, 1990.
  • Hyde, Lewis. Common as Air: Revolution, Art, and Ownership. Farrar, Straus and Giroux, 2010.

On The IndieWeb §

  • IndieWeb Wiki. https://indieweb.org/
  • Çelik, Tantek. “Own Your Data.” https://tantek.com/2020/015/t1/own-your-data
  • Winer, Dave. “Still Trying to Save the World.” http://scripting.com/

Primary Sources §

  • Mastodon. “What is Mastodon?” https://joinmastodon.org/
  • Matrix. “Matrix FAQ.” https://matrix.org/faq/
  • ENS Documentation. https://docs.ens.domains/
  • IPFS Docs. https://docs.ipfs.tech/

End of Chapter 4

Next: Chapter 5 — Triage Methodology: The Custodial Filter and Ethical Preservation

Part I • Theoretical Foundations & Taxonomy

Chapter 5: Triage Methodology

The Custodial Filter and Ethical Preservation

24 min read 5,072 words

Opening: The Impossible Choice §

October 16, 2016. Vine announces it will shut down in three months. Archive Team mobilizes immediately, but the math is brutal:

  • 200 million videos exist on Vine
  • Three months until shutdown
  • Limited volunteers, storage, and bandwidth

Even working around the clock, they can’t save everything. They must choose.

Do they prioritize:

  • Viral videos (most cultural impact, but already widely copied)?
  • Marginalized creators (underrepresented voices, but lower view counts)?
  • Complete user archives (preserving entire creator portfolios, but means fewer total creators saved)?
  • Representative sampling (cross-section of Vine culture, but many individual voices lost)?

Every choice means something else dies. Every video saved means another left behind.

This is triage—borrowed from battlefield medicine, where doctors must decide which wounded soldiers to treat first when resources are scarce. In emergency rooms, triage saves lives by allocating attention efficiently. In digital preservation, triage saves culture by allocating effort strategically.

But triage is agony. It forces us to confront uncomfortable truths:

  • Not everything can be saved
  • Some artifacts matter more than others
  • Scarcity requires hierarchy
  • Every preservation decision is also a decision to let something die

This chapter explores how to make those impossible choices—not perfectly (perfection is impossible), but ethically, systematically, and transparently.

We call this framework the Custodial Filter: a methodology for deciding what to preserve, when to preserve it, and when—painfully—to let go.


Part I: The Ethics of Triage §

Why Triage Is Necessary §

Infinite Culture, Finite Resources

The internet produces content at a rate no human effort can fully capture:

  • Twitter: 500 million tweets per day (2023)
  • YouTube: 720,000 hours of video uploaded daily
  • Instagram: 95 million photos and videos daily
  • TikTok: Unknown, but comparable to YouTube
  • Plus: Blogs, forums, Discord servers, newsletters, personal websites, etc.

Even with unlimited storage (which doesn’t exist), the labor of curation—adding metadata, providing context, ensuring accessibility—is scarce.

Platform Death Accelerates Urgency

When a platform announces shutdown, the timeline collapses:

  • GeoCities: 3 weeks warning
  • Vine: 3 months warning
  • Google Reader: 4 months warning
  • Tumblr NSFW purge: 2 weeks warning

In crisis mode, triage becomes life-or-death for artifacts.

Preservation Requires Stewardship

Saving bits is relatively cheap (storage costs drop constantly). But meaningful preservation requires:

  • Metadata creation (who, what, when, why, context)
  • Format migration (as technology evolves)
  • Access infrastructure (search, browse, display)
  • Legal navigation (copyright, privacy, consent)
  • Institutional maintenance (organizations must survive decades)

These activities consume human time and expertise—resources that will always be scarce.

The Ethical Stakes of Triage §

Who Decides What’s Worth Saving?

Triage decisions encode power and values:

  • If we prioritize “viral” content, we amplify mainstream voices and erase margins
  • If we prioritize “cultural significance,” we risk bias toward dominant cultures
  • If we prioritize ease of preservation, we lose complex, fragile artifacts
  • If we prioritize consent, we may lose important historical evidence

Every triage framework embodies ethical commitments, whether explicit or not.

The Permanence of Loss

Physical artifacts can be rediscovered—buried ruins excavated, manuscripts found in attics. But digital artifacts vanish completely when platforms shut down. There’s no archaeological dig 100 years later to recover what we failed to save.

Triage decisions are irreversible. What we don’t preserve now is lost forever.

The Burden of Custodianship

To preserve is to claim custodial responsibility:

  • You decide what future generations can know about this era
  • You become a gatekeeper—your choices shape historical memory
  • You bear ethical weight of what you saved and what you didn’t

This burden can’t be escaped. Even choosing not to preserve is a choice with consequences.


Part II: The Custodial Filter — A Five-Question Framework §

The Custodial Filter is a systematic methodology for triage. Before preserving any artifact, ask five questions:

Question 1: Cultural Significance §

Does this artifact represent a community, movement, or cultural moment that would otherwise be lost?

Criteria:

  • Representational value: Does it document an underrepresented community?
  • Historical importance: Does it capture a significant event or movement?
  • Aesthetic innovation: Does it represent creative achievement or technical pioneering?
  • Community meaning: Do people who created/used this consider it important?

High Significance Examples:

  • Early Black Twitter threads (2010-2015): Document emergence of hashtag activism (#BlackLivesMatter, #SayHerName)
  • Early trans YouTubers (2006-2012): Chronicle transition vlogs before mainstream visibility
  • GeoCities fan communities (1995-2000): Archive of early fandom, particularly marginalized fandoms (slash fiction, queer representation)

Lower Significance Examples:

  • Corporate spam accounts: Minimal cultural value, widely preserved elsewhere if needed
  • Duplicate viral videos: Already archived by multiple sources
  • Auto-generated content: Bot posts with no human creative input

Challenge: Whose Significance?

What’s “significant” is contested:

  • Academic historians prioritize different artifacts than community members
  • Mainstream culture dismisses subcultures as trivial (but those subcultures have rich internal meaning)
  • Future generations may value what present dismisses

Best Practice: Default to over-preservation when significance is uncertain. We can’t predict what future scholars will want to study.

Question 2: Technical Fragility §

How close to disappearance is this artifact?

Fragility Spectrum:

Critical (Hours/Days)

  • Platform announced shutdown imminent
  • Server errors suggest infrastructure collapse
  • DMCA takedowns being issued
  • Legal threats to hosting

High (Weeks/Months)

  • Platform announced future shutdown
  • Company in financial distress
  • Terms of Service changes pending (mass deletions coming)
  • Migration waves beginning (users leaving)

Medium (Years)

  • Platform declining but stable
  • No imminent shutdown threat
  • Content still accessible but endangered long-term

Low (Decades)

  • Stable institutions (library collections, government archives)
  • Already preserved with redundancy
  • Open formats, no proprietary lock-in

Triage Priority:

  • Critical fragility → Act immediately (even if cultural significance is uncertain)
  • Low fragility → Defer (focus on more endangered artifacts)

Example: GeoCities vs. Library of Congress

When both GeoCities and LOC’s web archive need attention:

  • GeoCities: Critical fragility (3 weeks to shutdown) → Priority 1
  • LOC: Low fragility (institutional stability, funded mandate) → Priority 3

Question 3: Rescue Difficulty §

How hard is this artifact to preserve?

Ease Assessment:

Easy (Can automate)

  • Static HTML pages (wget scraper)
  • Public APIs with bulk export
  • Standard formats (plain text, images, HTML)
  • Already-crawled by Internet Archive

Medium (Requires manual effort)

  • Dynamic content (JavaScript-heavy sites)
  • Private/login-walled content
  • Embedded media (Flash, Java applets)
  • Metadata extraction needed

Hard (Technical barriers)

  • Complex databases without export tools
  • DRM-protected content
  • Real-time/ephemeral content (Snapchat stories, Clubhouse rooms)
  • Server-side logic required for functionality

Very Hard (Near-impossible)

  • Fully encrypted with lost keys
  • Proprietary formats with no documentation
  • Deleted content with no backups
  • Hardware-specific content (arcade games requiring original boards)

Triage Tension:

Should you spend 100 hours preserving one hard artifact, or preserve 100 easy artifacts in the same time?

No universal answer, but factors to consider:

  • If hard artifact is uniquely significant (e.g., only documentation of a marginalized community), worth the effort
  • If easy artifacts are low-significance duplicates, hard artifact may be better use of time
  • If you’re in crisis mode (imminent shutdown), prioritize quantity (easy artifacts)

Example: Flash Games

Flashpoint Project prioritized Flash games (medium-hard difficulty) because:

  • High cultural significance (entire generation’s childhood)
  • Critical fragility (Flash Player discontinued)
  • Doable difficulty (emulation possible with effort)

They chose one hard project over many easy ones—and succeeded.

Question 4: Existing Redundancy §

Is someone else already preserving this?

Check for Redundancy:

  • Internet Archive’s Wayback Machine: Has it been crawled?
  • Library of Congress: Do they have it? (Web archive, Twitter archive)
  • Other institutions: University archives, national libraries, museums
  • Community efforts: Fan archives, Discord channels, subreddit backups
  • Individual users: Have creators exported their own content?

Redundancy Matrix:

Situation Action
No one preserving Urgent priority (you might be the only chance)
One fragile preservation Valuable redundancy (create backup of backup)
Multiple stable institutions Lower priority (focus elsewhere unless you add unique value)
Already in Internet Archive + LOC + universities Deprioritize (unless you’re doing different kind of curation)

Exception: “Preserve Differently”

Even if something is archived, you might preserve it differently:

  • Internet Archive: Comprehensive but minimal metadata
  • Your project: Smaller sample with rich contextualization
  • Both add value

Example: Vine

Internet Archive scraped Vine comprehensively (quantity). Individual fans created curated collections (quality—“Best Vines 2013-2017”). Both were valuable.

Should we preserve this?

This is the hardest question—and the one most often skipped. Just because you can preserve something doesn’t mean you should.

Ethical Red Flags:

Privacy Violations

  • Personal information shared under expectation of ephemerality (Snapchat-style content)
  • Medical, financial, or intimate details
  • Children’s content (especially if they can’t consent now)
  • Location data that could enable stalking

Potential for Harm

  • Revenge porn or non-consensual intimate images
  • Doxxing (personal addresses, phone numbers)
  • Harassment campaigns
  • Misinformation that continues to cause harm

Contested Consent

  • Creator deleted content intentionally (wanted it forgotten)
  • Content was private or “friends-only” (context collapse if made public)
  • Platform TOS forbade scraping (legal gray area)

Cultural Sensitivity

  • Indigenous knowledge that communities want kept within community
  • Religious or spiritual content with access restrictions
  • Closed cultural practices not meant for outsiders

Trauma and Re-traumatization

  • 9/11 jumper photos (newsworthy but deeply painful)
  • Mass shooting livestreams
  • Graphic violence or suffering

The Ethical Tension:

Preservation often conflicts with privacy/consent:

  • Historian’s view: “Everything is historically important; preserve now, restrict access if needed”
  • Privacy advocate’s view: “People have a right to be forgotten; preservation without consent is violence”

No Easy Resolution, but principles to guide:

Principle 1: Minimize Harm

  • If preserving causes direct, immediate harm (endangers someone’s safety), don’t do it
  • Example: Don’t archive doxxing threads that reveal someone’s address

Principle 2: Respect Explicit Deletion

  • If a creator intentionally deleted something (not platform-deleted), presume they wanted it gone
  • Exception: Public figures, historical importance (politicians deleting compromising tweets)

Principle 3: Restrict Access When Appropriate

  • Preserve but don’t make public (researcher-only access, embargoes)
  • Example: Archive controversial forum but require IRB approval to access

Principle 4: Community Consultation

  • When preserving community-created content, ask the community
  • Example: Indigenous archives often require tribal consultation

Principle 5: Transparent Decision-Making

  • Document why you preserved or didn’t
  • Allow for appeals/reconsideration

Case Study: Tumblr NSFW Purge

In 2018, Tumblr banned all “adult content,” deleting millions of posts, many of which were:

  • LGBTQ+ identity exploration
  • Sex education resources
  • Art (nudes, erotic fiction)
  • Sex worker portfolios

Ethical Dilemma: Should archivists preserve purged content?

Arguments FOR:

  • Cultural/historical significance (LGBTQ+ history)
  • Censorship resistance (corporation shouldn’t decide what’s “obscene”)
  • Creators may have lost only copies

Arguments AGAINST:

  • Some creators wanted content ephemeral (chosen not to archive personally)
  • Adult content has complex consent issues (performers may not want redistribution)
  • Legal risks (some purged content may have been illegal, archivists don’t want liability)

What Actually Happened:

  • Some archivists saved portions (research access only)
  • Many creators self-archived (exported their own blogs)
  • Much was permanently lost (no comprehensive rescue)

Ethical Assessment:

  • No single right answer
  • Case-by-case determination based on consent signals, cultural value, harm potential

Part III: The Triage Decision Matrix §

Combine all five questions into a scoring system to prioritize artifacts systematically.

Scoring Framework (0-5 scale for each dimension) §

Cultural Significance (0 = spam, 5 = irreplaceable cultural artifact)

Technical Fragility (0 = stable/safe, 5 = will disappear in hours)

Rescue Feasibility (0 = impossible, 5 = trivial to preserve; inverted for priority)

Redundancy Gap (0 = many redundant copies, 5 = unique, no other preservation)

Ethical Clarity (0 = serious ethical problems, 5 = clearly ethical to preserve)

Example Triage Matrix: Vine Shutdown §

Artifact Significance Fragility Feasibility Redundancy Ethics Total Priority
Viral memes (already copied) 4 5 5 2 5 21 Medium
Small creator archives 5 5 4 5 5 24 High
Corporate brand accounts 2 5 5 1 5 18 Low
Private accounts 3 5 3 5 2 18 Low (ethics)
Representative sample 4 5 5 4 5 23 High

Priority Ranking:

  1. Small creator archives (24 points) — Unique voices, no other preservation, highly fragile
  2. Representative sample (23 points) — Cultural cross-section, high feasibility
  3. Viral memes (21 points) — Significant but already widely copied
  4. Private accounts (18 points) — Ethical concerns override other factors
  5. Corporate accounts (18 points) — Low cultural value

Triage in Action: Three Scenarios §

Scenario 1: Imminent Shutdown (48 hours)

Situation: Small forum announces shutdown in 2 days. 10,000 posts, no warning.

Triage Decision:

  • Significance: Medium (small community, but may be only documentation)
  • Fragility: Critical (48 hours)
  • Feasibility: Medium (need to scrape + login walls)
  • Redundancy: High (probably no one else saving)
  • Ethics: Medium (public forum, but check for privacy issues)

Action: Immediate scrape. Archive everything, sort out curation later. In crisis, preservation > perfection.

Method:

  1. Use wget or HTTrack to scrape visible content
  2. Ask community members for database dump (if possible)
  3. Archive now, curate later (when not under deadline)

Scenario 2: Declining Platform (1-2 years warning)

Situation: Google+ shutdown announced for 2019. Company gives 1 year notice.

Triage Decision:

  • Significance: Medium (smaller than Facebook/Twitter, but had communities)
  • Fragility: High (shutdown certain) but not immediate
  • Feasibility: Medium-high (Google provided data export tools)
  • Redundancy: Low (Google+ not widely archived)
  • Ethics: High (users had export options, most content public)

Action: Systematic preservation with community partnership

Method:

  1. Partner with Internet Archive for Wayback crawls
  2. Create guides for users to export their own data
  3. Identify high-value communities (e.g., Photography+ had professional communities)
  4. Curate sample collections (not everything, but representative)
  5. Take full year to do it right (not crisis mode)

Scenario 3: Ongoing Platform with Contested Content

Situation: Twitter still operational, but waves of account suspensions. Some suspended accounts have historically important content.

Triage Decision:

  • Significance: Varies (some accounts very significant, others not)
  • Fragility: Medium (accounts suspended but may be reinstated, or may be permanent)
  • Feasibility: High (if archived before suspension; impossible after)
  • Redundancy: Low (Twitter doesn’t preserve suspended accounts)
  • Ethics: Complex (some suspensions justified, some censorship)

Action: Selective proactive archiving with ethical review

Method:

  1. Identify accounts with high historical/cultural value (activists, journalists, politicians)
  2. Proactively archive (before suspension) using tools like Twitter Archiver
  3. For already-suspended: check if Internet Archive captured (Wayback Machine)
  4. Ethical case-by-case: Don’t archive hate groups, do archive wrongfully suspended activists
  5. Restrict access for contentious material (researcher-only)

Part IV: Practical Triage Workflows §

Workflow 1: Crisis Triage (Platform Shutdown Imminent) §

Phase 1: Assess (Hours 1-4)

  1. How much time until shutdown?
  2. How much content exists?
  3. Who else is archiving?
  4. What tools are available?

Phase 2: Mobilize (Hours 4-24)

  1. Recruit volunteers (Archive Team, Twitter, Reddit)
  2. Set up infrastructure (servers, storage, coordination)
  3. Divide labor (different people scrape different sections)

Phase 3: Execute (Remaining time)

  1. Quantity over quality: Save everything you can
  2. Metadata is secondary (just get the bits)
  3. Accept losses (you won’t get everything)

Phase 4: Post-Shutdown

  1. Consolidate scraped data
  2. Remove duplicates
  3. Begin curation (add metadata, organize)
  4. Make accessible (upload to Internet Archive, create search interface)

Example: GeoCities Rescue

  • 3 weeks warning → Archive Team scraped 650GB
  • Post-shutdown → Organized into browseable torrent
  • Years later → Cameron’s World and other curated projects emerged

Workflow 2: Anticipatory Preservation (Platform Declining) §

Phase 1: Monitor (Ongoing)

  • Watch for signs of platform instability (layoffs, financial trouble, user exodus)
  • Begin proactive archiving before shutdown announced

Phase 2: Plan (When decline evident)

  1. Identify most valuable content (communities, creators, cultural artifacts)
  2. Assess redundancy (what’s already archived?)
  3. Develop curation strategy (can’t save everything, but can save representative sample)

Phase 3: Execute (Before crisis)

  1. Methodical crawling (not frantic scraping)
  2. Add metadata as you go
  3. Coordinate with platform (ask for data dumps, export tools)

Phase 4: Maintain (After shutdown)

  1. Preserve archives long-term (storage, format migration)
  2. Make accessible (search, browse, context)
  3. Document (write history of the platform for future scholars)

Example: LiveJournal Migration

  • Decline gradual (2007-2017)
  • Many users migrated to Dreamwidth, taking archives
  • Internet Archive captured public posts
  • By time Russian ownership happened (2017), most preservation already done

Workflow 3: Continuous Curation (Ongoing Platforms) §

Phase 1: Define Scope

  • You can’t archive the entire internet
  • Choose specific communities, topics, or creators to follow

Phase 2: Automate

  • Set up tools to continuously archive (RSS readers, auto-scrapers, bot accounts)
  • Example: ArchiveTeam’s “web sheriff” bots monitor for site deaths

Phase 3: Curate Regularly

  • Review captured content quarterly
  • Add metadata, context, interpretation
  • Identify gaps (what are you missing?)

Phase 4: Respond to Crises

  • When your monitored platforms face threats, escalate to crisis mode
  • You have head start (already archiving proactively)

Example: Internet Archive’s Wayback Machine

  • Continuous crawling since 1996
  • 800+ billion pages captured
  • When site dies, already have historical snapshots

Part V: Ethical Edge Cases §

Edge Case 1: The Deleted Tweet from a Public Figure §

Scenario: A politician tweets something racist, then deletes it 20 minutes later. Should you archive it?

Ethical Considerations:

FOR Archiving:

  • Public figure’s public statement (not private communication)
  • Accountability: politicians should be held responsible for their words
  • Historical record: deletion is an act of historical revisionism

AGAINST Archiving:

  • Person deleted it (signal they regret it, want it forgotten)
  • Could be taken out of context or misunderstood
  • Perpetuates harm by keeping racist content circulating

Custodial Filter Analysis:

  • Significance: High (public accountability)
  • Fragility: Critical (already deleted, may vanish from screenshots)
  • Feasibility: Easy (single tweet, text)
  • Redundancy: Medium (others likely screenshotted, but could be lost)
  • Ethics: Medium-high (public figure, accountability trumps right to be forgotten)

Recommendation: Preserve with context

  • Archive the tweet + surrounding context (what prompted it, reactions, apology if any)
  • Include in politician’s archival record
  • Make accessible (not hidden, but not amplified—no need to splash it on front page)

Edge Case 2: The Fan Fiction Archive §

Scenario: A LiveJournal community for a specific fandom (slash fiction, LGBTQ+ content) is abandoned. Creators have scattered. Should you archive?

Ethical Considerations:

FOR Archiving:

  • LGBTQ+ cultural history (much early queer culture happened in fandom)
  • Risk of permanent loss (creators may not have backups)
  • Literary/cultural value (transformative works, creative community)

AGAINST Archiving:

  • Many authors used pseudonyms, may not want real identities connected
  • Some authors were minors when writing (consent issues)
  • Fan fiction culture values ephemerality (archives disrupt gift economy)
  • Copyright gray area (transformative works, but still derivative)

Custodial Filter Analysis:

  • Significance: High (queer history, literary culture)
  • Fragility: High (no one maintaining it)
  • Feasibility: Medium (may need login, scraping fanfic sites common)
  • Redundancy: Low (likely not preserved elsewhere)
  • Ethics: Complex (consent unclear, cultural sensitivity needed)

Recommendation: Archive with restrictions

  1. Scrape the content (preserve the bits)
  2. Don’t make fully public (no Google indexing)
  3. Researcher access only (require application, explain use)
  4. Allow author-requested takedowns (if someone says “please remove my fic,” do it)
  5. Document the community culture (not just stories, but context of why this mattered)

Edge Case 3: The Hate Forum §

Scenario: A white supremacist forum announces shutdown. It documents radicalization pathways and extremist organizing. Should you archive?

Ethical Considerations:

FOR Archiving:

  • Research value (understanding radicalization, deradicalization efforts need data)
  • Legal accountability (evidence of planned violence)
  • Historical record (documenting extremism is important, even/especially if ugly)

AGAINST Archiving:

  • Amplifies hate speech (giving platform to harmful ideology)
  • Could be used as recruitment tool (if archive is public)
  • Privacy of victims (hate content often targets individuals)
  • Moral complicity (by preserving, are you endorsing?)

Custodial Filter Analysis:

  • Significance: Medium-high (historical/research value, but harmful)
  • Fragility: High (extremist sites often shut down by hosts or law enforcement)
  • Feasibility: Medium (may require Tor, technical barriers)
  • Redundancy: Low (mainstream archives avoid extremist content)
  • Ethics: Low (serious concerns about harm)

Recommendation: Very restricted archive, if at all

Option A (Maximum Security):

  1. Archive for research only (no public access)
  2. Require IRB approval + academic credentials to access
  3. Redact personal information of victims
  4. Coordinate with law enforcement (if active threats)
  5. Provide to hate-monitoring organizations (ADL, SPLC)

Option B (Don’t Archive):

  • Some things should be lost
  • If research value is low and harm potential is high, destruction is ethical
  • Document that it existed (metadata, description) without preserving content itself

Most archivists choose Option A: Preserve but lock down tightly. History includes ugly things, and understanding extremism requires evidence.

Edge Case 4: The Private Message Leak §

Scenario: Someone leaks a trove of private Discord messages revealing corporate malfeasance. The messages are newsworthy but were shared under expectation of privacy. Should you archive?

Ethical Considerations:

FOR Archiving:

  • Public interest (corporate wrongdoing should be documented)
  • Whistleblower protection (if original leaker is endangered, redundant copies help)
  • Historical record (evidence of how corporations operate behind closed doors)

AGAINST Archiving:

  • Privacy violation (people wrote those messages expecting privacy)
  • Consent (participants didn’t agree to archiving)
  • Collateral damage (leaks often include innocent bystanders’ private info)

Custodial Filter Analysis:

  • Significance: High (public interest, accountability)
  • Fragility: Medium (leak may be taken down via DMCA, threats)
  • Feasibility: Easy (already leaked, just need to copy)
  • Redundancy: Medium (likely others saving, but could be suppressed)
  • Ethics: Low-medium (privacy violation vs. public interest)

Recommendation: Selective archive with redaction

  1. Archive the newsworthy messages (evidence of wrongdoing)
  2. Redact personal information of non-involved parties (people who just happened to be in the server)
  3. Remove sensitive personal details (even of wrongdoers—focus on the malfeasance, not their kids’ names)
  4. Make available to journalists and researchers
  5. Consider time embargo (publish now, release full archive in 10 years when people involved are less vulnerable)

Principle: Public interest can override privacy, but minimize collateral damage.


Part VI: Institutional Triage Policies §

Building a Triage Policy for Your Organization §

If you’re creating an archive, museum, or preservation institution, codify your triage principles:

Policy Components:

1. Mission Statement

  • What are you preserving and why?
  • Example: “We preserve LGBTQ+ digital culture to ensure queer history isn’t erased”

2. Significance Criteria

  • What makes something worth preserving in your collection?
  • Be specific: representational gaps, community value, historical importance

3. Ethical Red Lines

  • What will you NOT preserve, no matter what?
  • Examples: “We do not preserve non-consensual intimate images” or “We do not archive active doxxing campaigns”

4. Restricted Access Guidelines

  • Under what conditions do you restrict access?
  • Who can access restricted materials?

5. Takedown Process

  • How can people request removal of material?
  • What’s the review process?

6. Transparency Commitment

  • How do you document triage decisions?
  • Do you publish criteria publicly?

Example: The Internet Archive’s Policy (Simplified)

  • Mission: “Universal access to all knowledge”
  • Significance: Broad crawling (no strict curation—preserve as much as possible)
  • Ethics: Respect robots.txt (if site owner says “don’t crawl,” they don’t), DMCA takedowns honored
  • Access: Public by default, but allow author/site owner opt-out
  • Transparency: Public-facing form for takedown requests, documents policies on website

Example: A Hypothetical Trans Archive’s Policy

  • Mission: “Preserve trans people’s digital self-documentation and community organizing”
  • Significance: Prioritize trans creators, especially early/formative content (pre-2010), survival resources, community organizing
  • Ethics: Strong consent focus—reach out to creators when possible, honor deletion requests, never out people
  • Access: Public access for educational/research use, but some material (private forums, DMs) restricted to trans researchers only
  • Transparency: Advisory board of trans community members reviews contested triage decisions

Part VII: When to Let Go §

The Hardest Lesson: Accepting Loss §

Not everything can be saved. Sometimes, the ethical choice—or the practical choice—is to let something die.

When to Let Go:

1. Ethical Harm Outweighs Value

  • If preserving actively hurts people (doxxing, revenge porn), don’t do it

2. No Viable Path to Preservation

  • Some artifacts are technologically impossible to save (encrypted with lost keys, hardware-specific with no working hardware)

3. Resources Better Spent Elsewhere

  • If saving one low-value artifact means letting a high-value artifact die, let the low-value one go

4. Respecting Intentional Ephemerality

  • Some cultures and communities value impermanence (Snapchat culture, Buddhist sand mandalas)
  • Forcing permanence violates cultural values

The Grief of Triage

Letting artifacts die is painful. You’re choosing what future generations can never know. You’re accepting that some stories will be lost, some voices silenced, some memories erased.

This grief is unavoidable. The role of the Archaeobytologist includes mourning.

But: Grief that paralyzes is counterproductive. Mourn, then act. Save what you can. Document what you couldn’t save (at least record that it existed). Move forward.

The Triage Paradox:

The better you get at triage, the more aware you become of loss. Beginners think they can save everything. Experts know they can’t—and carry the weight of every choice.

This is the burden of custodianship.


Conclusion: Triage as Ethical Practice §

The Custodial Filter isn’t a formula—it’s a framework for ethical deliberation. It forces you to ask hard questions:

  • What makes this artifact matter?
  • How urgently endangered is it?
  • Can we realistically save it?
  • Are others already saving it?
  • Should we save it?

Every triage decision is an ethical act. You’re deciding what the future can know about the past. You’re allocating scarce resources (time, labor, storage, attention). You’re potentially overriding someone’s wishes (to be forgotten, to be private).

These decisions should be:

  • Systematic (not arbitrary or impulsive)
  • Transparent (document your reasoning)
  • Revisable (be willing to reconsider)
  • Humble (acknowledge you could be wrong)

The Custodial Filter provides structure for these decisions—not certainty, but rigorous ethical thinking.

In the next chapter, we’ll explore the boundaries of Archaeobytology as a discipline—how it differs from adjacent fields, what makes it distinct, and why it deserves recognition as its own domain of study.

But first, practice triage. Look at your own digital life. What would you save if you had 48 hours to archive everything? What would you let go? And how would you justify those choices?

The Custodial Filter begins with seeing your own values clearly.


Discussion Questions §

  1. Personal Triage: If your email account announced shutdown in 48 hours, what would you prioritize saving? Why? What would you let go?

  2. Ethical Boundaries: Where do you draw the line? What content should never be archived, even if historically significant?

  3. Competing Values: How do you balance (a) preserving everything for future research vs. (b) respecting privacy and consent?

  4. Bias and Representation: How can triage avoid reproducing systemic biases (racism, sexism, class privilege)? Is “objective” triage possible?

  5. Institutional vs. Individual: Should triage decisions be made by institutions (museums, archives) or individuals (you with your hard drive)? What are the pros/cons of each?

  6. Future Regret: Imagine it’s 2075. What digital culture from 2020s do you think future historians will wish we’d preserved but didn’t?


Exercise: Conduct a Triage Simulation §

Scenario: You have 72 hours and 1TB of storage to archive a dying platform before it shuts down. The platform has:

  • 50,000 user accounts
  • 5 million posts (text, images, videos)
  • 200 communities/groups
  • 10 years of history

You cannot save everything. Conduct triage.

Part 1: Define Your Values (300 words)

  • What’s your preservation mission?
  • What criteria matter most to you (representation, popularity, rarity, etc.)?

Part 2: Apply the Custodial Filter (500 words)

Create a triage matrix for these artifact types:

  1. Viral posts (high engagement, widely seen)
  2. Marginalized community content (LGBTQ+, disability, etc.)
  3. Long-form creative work (fiction, art, tutorials)
  4. Personal journaling/diaries
  5. Corporate/brand accounts

Score each on:

  • Cultural Significance (0-5)
  • Technical Fragility (0-5)
  • Rescue Feasibility (0-5)
  • Redundancy Gap (0-5)
  • Ethical Clarity (0-5)

Part 3: Make Decisions (500 words)

  • Given your 1TB limit, what do you save?
  • What do you deprioritize or leave behind?
  • How do you handle ethical dilemmas (private content, deleted posts)?

Part 4: Reflect (200 words)

  • How did it feel to make these choices?
  • What surprised you about your own values?
  • Would you make different choices under different constraints?

Further Reading §

On Triage and Preservation Ethics §

  • Caswell, Michelle. “Seeing Yourself in History: Community Archives and the Fight Against Symbolic Annihilation.” The Public Historian 36, no. 4 (2014): 26-37.
  • Flinn, Andrew. “Community Histories, Community Archives: Some Opportunities and Challenges.” Journal of the Society of Archivists 28, no. 2 (2007): 151-176.
  • Jimerson, Randall. Archives Power: Memory, Accountability, and Social Justice. SAA, 2009.
  • Nissenbaum, Helen. Privacy in Context: Technology, Policy, and the Integrity of Social Life. Stanford, 2009.
  • Solove, Daniel. Nothing to Hide: The False Tradeoff Between Privacy and Security. Yale, 2011.
  • Rosen, Jeffrey. “The Right to Be Forgotten.” Stanford Law Review Online 64 (2012): 88.

On Digital Preservation Methods §

  • Brügger, Niels, and Ralph Schroeder, eds. The Web as History. UCL Press, 2017.
  • Kirschenbaum, Matthew, et al. “Digital Materiality: Preserving Access to Computers as Complete Environments.” iPRES (2009).
  • Archives Team. “So You Want to Archive a Website.” https://wiki.archiveteam.org/

On Ethics of Difficult Knowledge §

  • Simon, Roger, et al. “Witness as Study: Attending to the Testimonies of Trauma, Memory, and Injustice.” Equity & Excellence in Education 38, no. 3 (2005): 191-198.
  • Caswell, Michelle. Urgent Archives: Enacting Liberatory Memory Work. Routledge, 2021.

End of Chapter 5

Next: Chapter 6 — Discipline Formation and Boundaries: Why Archaeobytology Needs to Exist

Part I • Theoretical Foundations & Taxonomy

Chapter 6: Discipline Formation and Boundaries

Why Archaeobytology Needs to Exist

22 min read 4,686 words

Opening: The Question No One Asks §

At academic conferences, when you introduce yourself as an Archaeobytologist, the response is always the same:

“That’s interesting! So… what is that, exactly?”

You explain: “I study murdered digital platforms, preserve their artifacts, and build alternatives that resist future murders.”

They nod politely. Then: “Oh, so you’re a digital historian?” Or: “Like a computer scientist?” Or: “Is that part of library science?”

And you have to say: “Sort of, but not really. It’s… something else.”

This is the problem. Archaeobytology doesn’t fit neatly into existing academic boxes. It’s not quite history, not quite computer science, not quite library science, not quite media studies. It draws from all of them but belongs to none.

This ambiguity has consequences:

  • No dedicated funding streams (NSF? NEH? Where do we apply?)
  • No tenure-track jobs (departments don’t know where to hire Archaeobytologists)
  • No professional societies (where do we gather?)
  • No canonical texts (what do students read?)
  • No clear legitimacy (is this a “real” field or just a hobby?)

This chapter argues: Archaeobytology deserves to exist as its own discipline, not as a subfield of something else. We need our own departments, journals, conferences, and professional pathways.

But first, we must understand: How do disciplines form? What makes a field distinct? And what must Archaeobytology do to achieve legitimacy?


Part I: How Disciplines Are Born §

The Social Construction of Knowledge §

Disciplines aren’t natural categories—they’re socially constructed. There’s no inherent reason why “sociology” and “anthropology” are separate fields, or why “computer science” split from “electrical engineering.”

Disciplines form through:

1. Intellectual Coherence

  • A shared set of questions, methods, and theories
  • Example: Economics studies “allocation of scarce resources”; psychology studies “mind and behavior”

2. Institutional Infrastructure

  • Departments, degree programs, journals, conferences
  • Example: American Sociological Association (founded 1905) legitimized sociology

3. Professional Pathways

  • Jobs for people trained in the discipline
  • Example: Clinical psychology created careers for PhDs outside academia

4. Boundary Work

  • Defining what the field IS and what it ISN’T
  • Example: Anthropology distinguishes itself from sociology (culture vs. social structure)

5. Canonical Texts and Founders

  • Works everyone in the field must read
  • Example: Durkheim’s Suicide for sociology; Kuhn’s Structure of Scientific Revolutions for science studies

6. External Recognition

  • Funding agencies, governments, and universities accept the field as legitimate
  • Example: NSF created “Science and Technology Studies” program in 1970s

Case Study 1: How Digital Humanities Became a Discipline §

Origins (1960s-1980s): “Humanities Computing”

  • Scholars using computers for text analysis
  • Seen as technical skill, not a discipline
  • No departments, scattered practitioners

Critical Mass (1990s-2000s)

  • Internet makes digital methods essential
  • Conferences emerge: ACH (1978), ADHO (2005)
  • Journals launch: Computers and the Humanities (1966), Digital Humanities Quarterly (2007)

Institutionalization (2010s)

  • Universities create DH centers (Stanford, UVA, CUNY, UCL)
  • Tenure-track jobs appear with “digital humanities” in title
  • Funding: NEH Office of Digital Humanities (2008)

Legitimacy Achieved (2020s)

  • DH is recognized field with professional society, journals, degree programs
  • Still marginal (few standalone departments), but no longer dismissed as “not real scholarship”

Timeline: ~50 years from scattered practice to institutional recognition

Lessons for Archaeobytology:

  • Institutionalization takes decades
  • Need visible infrastructure (journals, conferences, centers)
  • External funding helps (NEH, Mellon, etc.)
  • Still face “legitimacy crisis” even after establishing infrastructure

Case Study 2: How Data Science Exploded §

Origins (2000s): Industry Demand

  • Companies needed people to analyze big data
  • No academic discipline—hired statisticians, computer scientists, physicists

Rapid Formalization (2010s)

  • Term “data science” popularized (~2012)
  • Bootcamps emerge (Galvanize, General Assembly)
  • Universities create programs (UC Berkeley, NYU, Columbia)
  • Professional society: Data Science Association (2013)

Ubiquity (2020s)

  • Data science everywhere: academia, industry, government
  • Hundreds of degree programs
  • High salaries drive enrollment

Timeline: ~10 years from buzzword to ubiquitous discipline

Key Difference from DH:

  • Industry demand accelerated institutionalization
  • Money attracted universities (lucrative master’s programs)
  • Less intellectual coherence (still debated what data science “is”), but strong professional pathways

Lessons for Archaeobytology:

  • Industry demand speeds legitimation (but can corrupt mission)
  • Professional pathways matter (if students can get jobs, universities create programs)
  • Fast institutionalization possible (but rare)

Case Study 3: Science and Technology Studies (STS) §

Origins (1970s): Coalition of Disciplines

  • Historians of science + sociologists of knowledge + philosophers of technology
  • Shared interest: how science/tech shape society (and vice versa)

Boundary Struggles (1980s-1990s)

  • “Science Wars”: conflict with scientists who felt STS was anti-science
  • Internal debates: constructivism vs. realism, actor-network theory vs. critical theory

Stabilization (2000s-present)

  • Professional society: Society for Social Studies of Science (4S, founded 1975)
  • Journals: Social Studies of Science, Science, Technology & Human Values
  • Departments: MIT, Cornell, UC San Diego, York, etc.
  • Identity: Interdisciplinary but distinct

Timeline: ~40 years to stable institutional form

Lessons for Archaeobytology:

  • Interdisciplinary origins are common (we’re not weird for drawing from multiple fields)
  • Boundary struggles are normal (expect pushback from adjacent fields)
  • Coalitional politics help (build alliances with sympathetic scholars in history, CS, library science)

Part II: What Makes Archaeobytology Distinct? §

The Archipelago Problem §

Archaeobytology currently exists as scattered islands of practice:

  • Archive Team (guerrilla digital archiving)
  • Internet Archive (institutional preservation)
  • Digital historians (studying past platforms)
  • Media archaeologists (theorizing dead media)
  • Platform studies scholars (analyzing platform affordances)
  • Right-to-repair activists (fighting for user sovereignty)
  • IndieWeb advocates (building decentralized alternatives)

These practitioners rarely talk to each other. They publish in different venues, attend different conferences, use different vocabularies. They’re doing related work but don’t see themselves as part of a unified field.

Archaeobytology proposes: These scattered practices belong together. They share:

  1. Core Problem: Platform death and digital dispossession
  2. Dual Method: Preservation (Archive) + Creation (Anvil)
  3. Normative Commitment: Digital sovereignty (Three Pillars)
  4. Ethical Framework: Triage and the Custodial Filter

Boundary Work: What Archaeobytology Is NOT §

To define a discipline, you must say what it excludes. Here’s what Archaeobytology is NOT:

Digital History:

  • Studies the past using digital methods
  • Analyzes historical sources (digitized archives, born-digital records)
  • Primary goal: Historical understanding

Archaeobytology:

  • Intervenes to create a future past (rescues artifacts before they vanish)
  • Studies platforms as they’re dying (not just retrospectively)
  • Primary goal: Preservation + building alternatives

Relationship: Digital historians are Archaeobytology’s users. We preserve the artifacts they study. But we’re not doing history—we’re doing applied preservation and system design.

NOT Computer Science (Though Technical)

Computer Science:

  • Develops algorithms, systems, languages
  • Values: Efficiency, correctness, performance
  • Questions: “How do we build this?” “What’s the optimal solution?”

Archaeobytology:

  • Uses CS methods (web scraping, emulation, distributed systems) but as tools, not ends
  • Values: Cultural preservation, user sovereignty, ethical curation
  • Questions: “What should be saved?” “Who owns this?” “How do we prevent future murders?”

Relationship: We need CS skills, but our questions are humanistic and political, not purely technical.

NOT Library Science (Though Archival)

Library Science:

  • Manages collections, provides access
  • Expert in metadata, cataloging, preservation standards
  • Works within institutional frameworks (libraries, universities, governments)

Archaeobytology:

  • Often works outside institutions (guerrilla archiving, legal gray areas)
  • Preserves things institutions won’t touch (ephemeral platforms, contested content)
  • Builds alternative systems (not just stewarding existing ones)

Relationship: Librarians are allies. We respect their expertise. But we operate in spaces they can’t (scraping copyrighted content, rescuing platforms without permission).

NOT Media Archaeology (Though Theoretical)

Media Archaeology:

  • Excavates dead media to theorize technological change
  • Philosophical and interpretive (Foucault, Kittler, Ernst)
  • Retrospective analysis

Archaeobytology:

  • Proactive preservation (we don’t wait for media to die; we intervene)
  • Applied practice (we scrape, we build, we organize)
  • Prospective design (we forge alternatives)

Relationship: Media archaeology gives us theory. We give them preserved artifacts to theorize about. But our work is grounded in doing, not just thinking.

NOT Activism (Though Political)

Activism:

  • Mobilizes for policy change
  • Protest, advocacy, organizing
  • Values change over documentation

Archaeobytology:

  • Documents and builds (we create archives and tools, not just campaigns)
  • Scholarly methods (research, publication, teaching)
  • Values preservation alongside change

Relationship: Many Archaeobytologists are activists (fighting for right to archive, platform accountability). But activism alone isn’t Archaeobytology—we also do scholarship.

What Archaeobytology IS: A Synthetic Definition §

Archaeobytology is the study and practice of:

  1. Excavating digital artifacts endangered by platform death, obsolescence, or corporate murder
  2. Preserving those artifacts with technical fidelity and cultural context
  3. Curating collections that make sense of vast data, applying ethical triage
  4. Interpreting artifacts so future generations understand their significance
  5. Building tools, protocols, and institutions that embody digital sovereignty
  6. Advocating for laws and norms that protect digital culture from erasure
  7. Teaching others to do all of the above

Unique Combination:

  • Technical + humanistic
  • Retrospective (Archive) + prospective (Anvil)
  • Scholarly + activist
  • Individual practice + institutional design

No other field does all of this.


Part III: The Legitimacy Gap §

Why Archaeobytology Currently Lacks Legitimacy §

1. No Departments

  • You can’t get a PhD in Archaeobytology
  • Universities don’t hire “Archaeobytologists”
  • Students interested in this work must choose other departments (History, CS, Library Science, Media Studies)

2. No Dedicated Funding

  • NSF funds computer science (but we’re not CS)
  • NEH funds humanities (but we do technical work)
  • IMLS funds libraries (but we’re not traditional librarians)
  • We fall through cracks in funding taxonomies

3. No Professional Society

  • No “American Archaeobytological Association”
  • Practitioners scattered across multiple conferences (ADHO, SAA, 4S, ACM)
  • No unified community

4. No Canon

  • What books should every Archaeobytologist read?
  • Currently, reading lists are ad hoc (Kirschenbaum? Chun? Parikka? Doctorow? All of the above?)

5. No Clear Career Path

  • Where do you work after getting trained in Archaeobytology?
  • Internet Archive? Universities (but which department)? Tech companies (but doing what)?

6. Disciplinary Prejudice

  • Humanists see us as “too technical” (not real humanities)
  • Computer scientists see us as “not technical enough” (applied work, not theory)
  • Librarians see us as “reckless” (scraping without permission)
  • Activists see us as “too academic” (publishing papers instead of protesting)

We’re stuck in no-man’s-land between disciplines.

The Consequences of Illegitimacy §

For Students:

  • Can’t major in Archaeobytology (must choose proximate field, then specialize)
  • Dissertations get challenged (“Is this really History?” “Is this really CS?”)
  • Job market brutal (applying for jobs in History with “too much CS,” or vice versa)

For Practitioners:

  • Struggle to get tenure (unclear evaluation criteria)
  • Difficulty publishing (journals don’t know what to do with cross-disciplinary work)
  • Funding rejections (“This doesn’t fit our program”)

For the Field:

  • Slow growth (hard to recruit students if no clear pathway)
  • Knowledge fragmentation (practitioners don’t know what others are doing)
  • Lost opportunities (projects don’t happen because no institutional home)

For Society:

  • Platforms keep murdering culture (no unified opposition)
  • Preservation happens ad hoc (no systematic approach)
  • Alternatives fail to scale (no institutional support)

Part IV: Building Archaeobytology as a Discipline §

The Infrastructure We Need §

If Archaeobytology is to become legitimate, we need:

1. Knowledge Infrastructure

Journals:

  • Journal of Archaeobytology (peer-reviewed, interdisciplinary)
  • Publishes: Technical methods, ethical frameworks, case studies, theoretical essays, institutional designs

Conferences:

  • Annual Archaeobytology Conference (like ADHO for DH, or 4S for STS)
  • Brings together archivists, builders, scholars, activists
  • Creates community and shared identity

Textbooks and Handbooks:

  • This textbook is a start
  • Need: Handbook of Digital Preservation Methods
  • Need: Archaeobytological Theory: A Reader
  • Standardizes knowledge, creates canon

Online Platforms:

  • Archaeobytology Wiki (documenting methods, case studies, tools)
  • Forum for practitioners (discuss triage dilemmas, share technical solutions)
  • Repository of syllabi, assignments, datasets

Archives and Datasets:

  • Shared collections for teaching and research
  • Example: “The Murdered Platforms Database” (comprehensive data on every platform shutdown)

2. Institutional Anchors

University Programs:

  • Start with certificates and minors (“Certificate in Digital Preservation and Sovereignty”)
  • Grow to master’s programs (professional degree for archivists, curators)
  • Eventually: PhD programs (train next generation of scholars)

Centers and Institutes:

  • “Center for Digital Sovereignty” (like DH centers)
  • Provides: Servers for student projects, archival storage, research funding, speaker series

Labs:

  • “Preservation Lab” (students learn scraping, emulation, forensics)
  • “Anvil Lab” (students build protocols, tools, platforms)

Model: How Digital Humanities Did This

  • Stanford’s Center for Spatial and Textual Analysis (CESTA)
  • UVA’s Scholars’ Lab
  • CUNY’s GC Digital Initiatives
  • Start with grants, prove value, become permanent

3. Professional Pathways

Academic Track:

  • Tenure-track jobs in “Archaeobytology and Digital Culture”
  • Housed in: History depts, Media Studies, iSchools, or new Archaeobytology depts

Practitioner Track:

  • Archivist roles at Internet Archive, museums, libraries
  • “Digital Preservation Specialist” (job title that emphasizes Archaeobytology skills)

Industry Track:

  • Tech companies hiring “digital sovereignty architects”
  • Platform companies (ironically) needing people to design ethical data export/preservation

Non-Profit Track:

  • Working at EFF, Internet Archive, Creative Commons, Wikimedia
  • “Digital Rights Advocate” roles

Consulting:

  • Helping organizations design preservation strategies
  • Advising on platform alternatives (cooperatives, federated systems)

Certification:

  • “Certified Archaeobytologist” credential (like Certified Archivist)
  • Signals expertise to employers

4. External Recognition

Funding Programs:

  • NEH: “Archaeobytology Preservation Grants”
  • NSF: “Digital Sovereignty Infrastructure” program
  • Mellon Foundation: “Murdered Platform Archives” initiative

Government Acknowledgment:

  • Library of Congress hires Archaeobytologists
  • National Archives develops Archaeobytology methods
  • UNESCO recognizes digital culture preservation as essential

Public Visibility:

  • Popular books on Archaeobytology (like The Shallows for internet criticism)
  • Documentaries about platform death and preservation
  • Op-eds in NYT, Atlantic, Wired by Archaeobytologists

Part V: The 10-20 Year Roadmap §

Phase 1: Emergence (Years 1-5) — We Are Here §

Current State (2025):

  • Scattered practitioners doing Archaeobytology without calling it that
  • This textbook is one of first attempts to codify the field
  • No formal infrastructure (yet)

Goals for Phase 1:

  • Name the discipline: Get people to start calling themselves Archaeobytologists
  • Create online community: Wiki, forum, Discord/Slack for practitioners
  • First conference: Host “Archaeobytology 2026” (even if small—50 people)
  • First journal issue: Launch Journal of Archaeobytology (online, open access)
  • Secure initial grants: Mellon Foundation, NEH, Mozilla Foundation

Metrics of Success:

  • 100+ people identify as Archaeobytologists
  • 5-10 universities offer courses with “Archaeobytology” in title
  • 3-5 published articles citing Archaeobytology as a discipline

Phase 2: Coalition Building (Years 6-10) §

Goals for Phase 2:

  • Professional society: Found “Society for Archaeobytology” (or “Digital Sovereignty Studies”)
  • Grow conference: 200-300 attendees, international
  • Launch degree programs: First master’s in Archaeobytology (probably at iSchool or interdisciplinary program)
  • Establish centers: 3-5 universities have “Centers for Digital Sovereignty”
  • Policy advocacy: Archaeobytologists testify at hearings, draft model legislation

Metrics of Success:

  • 500+ self-identified Archaeobytologists
  • 20-30 universities teaching Archaeobytology courses
  • 10+ tenure-track jobs with “Archaeobytology” or “Digital Sovereignty” in description

Phase 3: Institutionalization (Years 11-15) §

Goals for Phase 3:

  • PhD programs: First dissertations in Archaeobytology
  • Textbook adoption: 50+ universities using this textbook or similar
  • Funding streams: NSF/NEH have dedicated Archaeobytology programs
  • Public recognition: NYT runs feature on “the Archaeobytologists saving the internet”

Metrics of Success:

  • 2,000+ Archaeobytologists
  • 50+ universities with programs (certificates, minors, concentrations)
  • 5+ PhD programs
  • Professional certification launched

Phase 4: Maturity (Years 16-20) §

Goals for Phase 4:

  • Standalone departments: First “Department of Archaeobytology and Digital Sovereignty” (like STS departments)
  • Canon established: Everyone agrees on core texts
  • Career pathways clear: Students know how to become Archaeobytologists
  • Public impact: Laws passed influenced by Archaeobytology research

Metrics of Success:

  • 5,000+ Archaeobytologists
  • 100+ universities with programs
  • 10+ standalone departments or institutes
  • Field is recognized by universities, funding agencies, governments

Timeline Reality Check:

  • Digital Humanities: ~50 years to current state (still marginal)
  • Data Science: ~10 years to ubiquity (but industry-driven)
  • STS: ~40 years to stable discipline

Realistic Expectation: 20-30 years to full legitimacy. But meaningful impact possible much sooner (5-10 years).


Part VI: Threats to Discipline Formation §

Threat 1: Disciplinary Capture §

Risk: Established fields absorb Archaeobytology as a subfield, preventing independence.

Scenarios:

  • History departments claim Archaeobytology as “digital history”
  • CS departments subsume it as “digital preservation” (purely technical)
  • Library schools treat it as “web archiving” (narrowly applied)

Consequence: Archaeobytology’s unique synthesis (Archive + Anvil, technical + humanistic, scholarly + activist) gets fragmented. Each discipline takes the parts they understand and discards the rest.

Defense:

  • Insist on synthetic identity: Archaeobytology is not reducible to any existing field
  • Build coalitions across disciplines (harder to capture if multiple fields claim us)
  • Create independent infrastructure (journal, conference, society) that isn’t controlled by existing disciplines

Threat 2: Industry Co-optation §

Risk: Tech companies see value in Archaeobytology, hire practitioners, dilute mission.

Scenarios:

  • Facebook hires “Digital Preservation Specialists” to archive deleted content (for ads/AI training)
  • Blockchain companies claim to be “Archaeobytologists” (conflating crypto speculation with sovereignty)
  • Platform companies use Archaeobytology rhetoric to greenwash extractive practices

Consequence: Field becomes associated with corporate interests, loses critical edge, alienates activist practitioners.

Defense:

  • Value clarity: Center the Three Pillars and anti-platform politics
  • Ethical standards: Professional code that excludes surveillance-capitalism work
  • Critical scholarship: Maintain academic critique of platforms (not just working for them)

Threat 3: Internal Fragmentation §

Risk: Practitioners can’t agree on boundaries, methods, or values. Field splinters.

Scenarios:

  • “Radical Archaeobytologists” (activists) vs. “Academic Archaeobytologists” (scholars) split
  • Methodological wars: “True preservation requires bit-perfect forensics” vs. “Triage means accepting good-enough captures”
  • Ethical divides: “Archive everything” vs. “Consent above all”

Consequence: No unified identity, infrastructure fails, discipline never gels.

Defense:

  • Big tent: Accommodate methodological diversity (multiple approaches valid)
  • Core values: Agree on essentials (Three Pillars, Custodial Filter) while debating details
  • Generosity: Don’t excommunicate people for disagreements (pluralism is strength)

Threat 4: Funding Droughts §

Risk: Foundations and agencies don’t fund Archaeobytology; infrastructure collapses.

Scenarios:

  • Economic recession cuts humanities/tech funding
  • Political shifts defund preservation and digital rights
  • Competing priorities (AI, climate) absorb available grants

Consequence: Can’t pay for journals, conferences, centers. Practitioners leave for funded fields.

Defense:

  • Diversify funding: Multiple sources (government, foundations, individual donations, earned revenue)
  • Demonstrate impact: Show that Archaeobytology work matters (saves culture, influences policy, creates economic value)
  • Partnerships: Work with established institutions (libraries, museums) that have stable funding

Threat 5: Irrelevance §

Risk: Platforms stop dying (monopolies stabilize), or new preservation technologies make Archaeobytology obsolete.

Scenarios:

  • Governments regulate platforms, mandate data portability, fund public archives → crisis solved, Archaeobytology not needed
  • Blockchain/IPFS “solves” preservation → technical solution makes human curation irrelevant
  • Platforms become permanent monopolies, shutdowns stop → no more murders to document

Consequence: Field loses urgency, students don’t enroll, discipline fades.

Reality Check: This threat is unlikely. Platform death will continue. New technologies create new preservation challenges. Human curation will always be needed.

Defense:

  • Adaptive mission: If some problems are solved (great!), focus on remaining ones
  • Expansive definition: Archaeobytology isn’t just about shutdowns—it’s about sovereignty, curation, interpretation (always needed)

Part VII: Adjacent Disciplines as Allies §

Archaeobytology doesn’t need to fight existing fields—it can collaborate:

Digital Humanities §

  • We offer: Preserved digital artifacts for their historical research
  • They offer: Methodological expertise (text mining, network analysis, visualization)
  • Collaboration: Joint projects analyzing murdered platforms

Computer Science §

  • We offer: Real-world problems needing technical solutions (emulation, distributed storage, protocol design)
  • They offer: Engineering expertise
  • Collaboration: CS students build tools for Archaeobytology projects (win-win)

Library and Information Science §

  • We offer: Knowledge of endangered digital content and preservation urgency
  • They offer: Metadata standards, long-term stewardship, institutional partnerships
  • Collaboration: Librarians curate what we rescue

Science and Technology Studies §

  • We offer: Case studies of platform power, technological politics
  • They offer: Theoretical frameworks (actor-network theory, social construction of technology)
  • Collaboration: STS scholars theorize; we provide empirical grounding

Media Studies §

  • We offer: Preservation of media objects for analysis
  • They offer: Cultural analysis, critical theory
  • Collaboration: Joint teaching (they analyze media; we preserve it)

Law and Policy §

  • We offer: Evidence of platform harms and preservation needs
  • They offer: Legal expertise (copyright, privacy, platform regulation)
  • Collaboration: Draft legislation for right to archive, data portability

Strategy: Be a boundary organization—work across disciplines while maintaining distinct identity.


Part VIII: What You Can Do Right Now §

Whether you’re a student, practitioner, or professor, you can help build Archaeobytology:

If You’re a Student §

1. Call Yourself an Archaeobytologist

  • In your bio, on your CV, in conversations
  • Naming creates identity

2. Propose Courses

  • Ask your department to offer “Introduction to Archaeobytology”
  • Use this textbook

3. Write Your Thesis on It

  • Dissertations/theses create scholarly legitimacy
  • Cite Archaeobytology as your field

4. Join the Community

  • Find others doing this work (Twitter, Discord, conferences)
  • Build networks

If You’re a Practitioner §

1. Publish Your Work

  • Write about your preservation projects
  • Document methods (tutorials, case studies)
  • Contribute to building canon

2. Attend/Organize Conferences

  • Present at existing venues (ADHO, SAA, 4S)
  • Organize Archaeobytology sessions or workshops
  • Eventually: Host Archaeobytology Conference

3. Seek Funding

  • Apply for grants explicitly for “Archaeobytology research”
  • Force funding agencies to engage with the term

4. Mentor Students

  • Train next generation
  • Create clear pathways

If You’re a Professor §

1. Teach Archaeobytology Courses

  • Offer courses with “Archaeobytology” in title
  • Adopt this textbook

2. Hire Archaeobytologists

  • When your department has an opening, advocate for “Archaeobytology specialization”
  • Write job ads that name the field

3. Start a Center

  • Apply for grants to create “Center for Digital Sovereignty”
  • Provide institutional home

4. Publish Research

  • Cite Archaeobytology explicitly in your work
  • Build scholarly community

If You’re an Administrator §

1. Create Programs

  • Certificate, minor, or master’s in Archaeobytology
  • Proves demand, attracts students

2. Support Infrastructure

  • Fund journals, conferences, speaker series
  • Provide space and resources

3. Hire Faculty

  • Create positions in Archaeobytology
  • Show universities this is a legitimate field

Conclusion: The Discipline That Must Exist §

Archaeobytology exists because it must. The forces that murder digital culture—platform capitalism, surveillance economics, planned obsolescence—are accelerating. We need a discipline dedicated to:

  • Preserving what platforms murder
  • Building alternatives that resist murder
  • Training people to do both
  • Advocating for laws that protect digital sovereignty

No existing field does this comprehensively. Each adjacent discipline handles part of the problem, but no one takes responsibility for the whole.

Archaeobytology fills this gap.

We’re not trying to replace History, Computer Science, or Library Science. We’re trying to create a home for work that falls between them—work that’s too technical for humanists, too humanistic for engineers, too radical for institutions, and too scholarly for activists.

This textbook is a founding document. By reading it, teaching from it, citing it, and building on it, you’re helping create the discipline.

In 20 years, there might be Archaeobytology departments at major universities. Students might major in it. There might be thousands of practitioners. Laws might protect digital culture because Archaeobytologists advocated for them.

Or this might remain a marginal practice, known only to specialists.

That depends on us. Disciplines don’t form spontaneously—they’re built through collective effort. By calling ourselves Archaeobytologists, teaching Archaeobytology, funding Archaeobytology, and practicing Archaeobytology, we make the discipline real.

In the next chapter, we begin Part II: Excavation and Forensics. Now that we understand what Archaeobytology is and why it needs to exist, we’ll learn how to do it—starting with the methods for excavating digital artifacts before they vanish.

The theory is complete. Now the practice begins.


Discussion Questions §

  1. Disciplinary Identity: Do you consider yourself an Archaeobytologist? If not, what field do you identify with? If yes, when did you adopt that identity?

  2. Boundary Work: Should Archaeobytology be a discipline, or a subfield of something else? What would we gain/lose by remaining interdisciplinary?

  3. Legitimacy Politics: What would it take for universities to recognize Archaeobytology as legitimate? Is academic legitimacy even desirable (or does it risk co-optation)?

  4. Career Pathways: If you wanted a career in Archaeobytology, what would your path look like? What obstacles would you face?

  5. Threat Assessment: Which threat to discipline formation (capture, co-optation, fragmentation, funding drought, irrelevance) seems most serious? How would you defend against it?

  6. Personal Action: What’s one concrete thing you could do in the next month to help build Archaeobytology as a discipline?


Exercise: Design Your Dream Archaeobytology Program §

Task: You’ve been hired to create the world’s first Archaeobytology program at a university. Design it.

Part 1: Program Structure (500 words)

  • What degree(s)? (Certificate, minor, BA, MA, PhD?)
  • What department(s) house it? (New department, or joint program?)
  • How many courses? What’s the curriculum?

Part 2: Sample Syllabus (1000 words)

Create a syllabus for one course:

  • “Introduction to Archaeobytology” (undergraduate survey)
  • OR “Advanced Triage and Preservation” (graduate seminar)
  • OR “Building Sovereign Systems” (technical course)

Include:

  • Learning objectives
  • Weekly topics
  • Readings (5-10 per week)
  • Assignments
  • How this course fits in larger program

Part 3: Institutional Infrastructure (500 words)

  • What facilities/resources do students need? (Servers, storage, lab space?)
  • What partnerships? (Internet Archive, local libraries, tech companies?)
  • How do you fund it? (Grants, tuition, endowment?)

Part 4: Career Pathways (500 words)

  • What jobs can graduates get?
  • How do you help them find employment?
  • What skills make them competitive?

Part 5: Reflection (300 words)

  • What’s the biggest challenge to launching this program?
  • How do you convince your university to approve it?
  • Would you want to be a student in this program? Why/why not?

Further Reading §

On Discipline Formation §

  • Abbott, Andrew. Chaos of Disciplines. University of Chicago Press, 2001.
  • How academic disciplines form, fragment, and compete

  • Klein, Julie Thompson. Interdisciplining Digital Humanities. University of Michigan Press, 2015.

  • Case study of DH’s struggle for disciplinary legitimacy

  • Kuhn, Thomas. The Structure of Scientific Revolutions. University of Chicago Press, 1962.

  • Classic on paradigm shifts and scientific disciplines (though focused on natural sciences)

  • Small, Mario Luis. “How to Conduct a Mixed Methods Study: Recent Trends in a Rapidly Growing Literature.” Annual Review of Sociology 37 (2011): 57-86.

  • On methodological pluralism in new fields

On Boundary Work §

  • Gieryn, Thomas. “Boundary-Work and the Demarcation of Science from Non-Science.” American Sociological Review 48, no. 6 (1983): 781-795.
  • Classic on how disciplines define themselves by exclusion

  • Star, Susan Leigh, and James Griesemer. “Institutional Ecology, ‘Translations’ and Boundary Objects.” Social Studies of Science 19, no. 3 (1989): 387-420.

  • How interdisciplinary work creates “boundary objects” (like Archaeobytology itself)

On Academic Legitimacy §

  • Burawoy, Michael. “For Public Sociology.” American Sociological Review 70, no. 1 (2005): 4-28.
  • On scholarship engaging public, not just academy (relevant to Archaeobytology’s activist dimension)

  • Posner, Miriam. “Here and There: Creating DH Community.” In Debates in the Digital Humanities 2016, edited by Matthew Gold and Lauren Klein, 265-276. University of Minnesota Press, 2016.

  • Building scholarly community in interdisciplinary field

On Professional Pathways §

  • Nowviskie, Bethany. “On the Origin of ‘Hack’ and ‘Yack.’” In Debates in the Digital Humanities, edited by Matthew Gold, 66-73. University of Minnesota Press, 2012.
  • On tension between doing (hacking) and talking (yacking) in DH (relevant to Archaeobytology)

  • Rockwell, Geoffrey, and Stéfan Sinclair. Hermeneutica: Computer-Assisted Interpretation in the Humanities. MIT Press, 2016.

  • On building scholarly careers in computational humanities

End of Chapter 6 — End of Part I: Foundations

Next: Part II — Excavation & Forensics Chapter 7 — Archaeological Methods for Digital Artifacts

Part II • Excavation Methods & Digital Forensics

Chapter 7: Archaeological Methods for Digital Artifacts

19 min read 4,109 words

Opening: The Dig Site Is Ephemeral §

In 1922, Howard Carter discovered Tutankhamun’s tomb. The artifacts had been buried for 3,000 years. They would remain buried for 3,000 more if Carter didn’t act. But once found, he had time—years to carefully excavate, photograph, catalog, and preserve each object.

In 2009, Archive Team discovered that GeoCities would shut down in three weeks. The “artifacts” had existed for 15 years. They would exist for 21 more days, then vanish forever. No time for careful documentation. No room for archaeological precision. Just frantic scraping before the servers went dark.

This is the fundamental difference between physical and digital archaeology:

Physical archaeology:

  • Sites persist for centuries
  • Excavation is slow, methodical, non-destructive
  • You can return to a site years later

Digital archaeology:

  • Sites vanish in days or weeks
  • Excavation is fast, opportunistic, often destructive (scraping overloads servers)
  • You get one chance—once the platform dies, it’s gone

Yet despite these differences, physical archaeology offers valuable methods for digital practice. Stratigraphic analysis, site surveys, provenance tracking, and ethical excavation frameworks all translate to digital contexts—if adapted properly.

This chapter explores how to excavate digital artifacts using archaeological methods modified for digital ephemera. You’ll learn:

  • Site reconnaissance and mapping
  • Stratigraphic analysis of digital layers
  • Excavation techniques (scraping, API harvesting, forensic recovery)
  • Provenance and chain-of-custody documentation
  • Ethical frameworks for excavation

By the end, you’ll know how to approach a dying platform like an archaeological dig site—systematic, ethical, and effective.


Part I: Site Reconnaissance — Mapping the Digital Landscape §

Before You Dig: Understanding the Site §

Physical archaeologists don’t start digging randomly. They survey the site, create maps, test soil composition, and plan their excavation strategy. Digital archaeologists must do the same.

Step 1: Platform Architecture Assessment §

Goal: Understand the platform’s technical structure before attempting to preserve it.

Questions to Answer:

1. What type of platform is this?

  • Static website (HTML/CSS, easy to scrape)
  • Dynamic web app (JavaScript-heavy, requires browser automation)
  • Mobile app (API-based, may require reverse engineering)
  • Forum/BBS (database-driven, need to capture structure)
  • Social network (graph-based, relationships matter as much as content)

2. What are the data types?

  • Text (posts, comments, messages)
  • Images (user photos, avatars, UI elements)
  • Videos (hosted on platform or embedded from elsewhere?)
  • Metadata (timestamps, user IDs, like counts, tags)
  • Relationships (follows, friends, replies, shares)

3. What’s the scale?

  • How many users?
  • How much content (posts, pages, files)?
  • How much storage required?

4. What are the access patterns?

  • Public (anyone can view)
  • Login-required (need account)
  • Private/friends-only (restricted access)
  • Ephemeral (content disappears after viewing, like Snapchat)

5. What are the technical barriers?

  • Rate limiting (how many requests per hour?)
  • JavaScript rendering (can’t scrape with simple wget)
  • CAPTCHAs (need human intervention)
  • DRM/encryption (legally/technically protected)
  • APIs (do they exist? are they documented?)

Example: GeoCities Architecture Assessment (2009)

Dimension Assessment
Type Static HTML sites (mostly)
Data types HTML, images, GIFs, MIDI files, JavaScript
Scale ~30 million sites, estimated TB of data
Access Public (no login required)
Barriers Rate limiting (Yahoo would block aggressive scrapers), broken links (sites linked to each other, many links dead)
Strategy Distributed scraping (many volunteers, different IPs), prioritize unique content over duplicates

Step 2: Existing Documentation §

Check what’s already known:

Internet Archive’s Wayback Machine:

  • Has it been crawled? When? How comprehensively?
  • Gaps in coverage?

Platform’s Official Archives:

  • Does the platform provide data export tools?
  • What format? How complete?

Community Knowledge:

  • Are there fan wikis, documentation, user guides?
  • Former employees willing to share insider knowledge?

Technical Documentation:

  • API documentation (if APIs exist)
  • Terms of Service (what’s legal to scrape?)
  • robots.txt (what does platform want crawled?)

Example: Vine Documentation Check (2016)

  • Wayback Machine: Some Vines captured, but incomplete (many videos not archived)
  • Official Export: Vine provided “Download Your Vines” tool (good, but required users to act)
  • Community: Vine Wiki documented popular creators, memes, culture
  • API: Public API existed (allowed bulk downloading until shutdown)

Decision: Use API while it exists, supplement with manual scraping for videos API misses.

Step 3: Reconnaissance Scraping §

Goal: Capture a small sample to understand structure before full excavation.

Method:

  1. Scrape 100-1000 pages/posts (small representative sample)
  2. Analyze structure:
  3. What HTML tags are used?
  4. Where is metadata stored? (JSON in page source? Separate API calls?)
  5. What’s the URL pattern? (Can you enumerate all pages?)
  6. Test tools:
  7. Does wget work? Or need Selenium (browser automation)?
  8. How fast can you scrape without getting blocked?

Deliverable: Reconnaissance report documenting:

  • Platform structure
  • Technical barriers
  • Estimated scale
  • Recommended tools
  • Estimated time to complete full excavation

Example: Small Forum Reconnaissance

bash
Platform: Example Forum (phpBB)
Date: 2024-11-15
Estimated Shutdown: 2024-12-01 (15 days)

Structure:

- Forum software: phpBB 3.2
- Content: 10,000 threads, ~50,000 posts
- Users: 1,500 registered, ~500 active

Technical Assessment:

- Public access (no login for reading)
- Standard HTML structure (easy to parse)
- URL pattern: /viewtopic.php?t=[thread_id]
- Thread IDs appear sequential (can enumerate)

Barriers:

- Rate limit: ~60 requests/minute before 503 errors
- Some images externally hosted (may be lost)
- User profile pages require login (skip for now)

Strategy:

- Use wget with --wait=1 (stay under rate limit)
- Scrape all threads over 5 days
- Capture HTML + images
- Parse HTML to extract structured data (JSON)

Estimated Storage: 2-5GB
Estimated Time: 5 days (continuous scraping)

Part II: Stratigraphic Analysis — Understanding Digital Layers §

Physical archaeologists use stratigraphy—the study of layers—to understand how a site was formed over time. Lower layers are older; upper layers are more recent. Disruptions in layers indicate events (fires, floods, invasions).

Digital platforms also have layers. Understanding them is crucial for preservation.

Digital Stratigraphy: The Technology Stack §

Layer 1: Content (Surface Layer)

  • What users see: posts, images, videos
  • This is the “archaeological treasure”—the artifacts themselves

Layer 2: Metadata (Context Layer)

  • Timestamps, user IDs, like counts, tags, geolocation
  • Essential for interpreting content

Layer 3: Relationships (Social Layer)

  • Follower graphs, reply threads, retweets/shares
  • Network structure that gives content meaning

Layer 4: Platform Affordances (Infrastructure Layer)

  • Character limits (Twitter’s 140/280), video length limits (Vine’s 6 seconds)
  • UI design (Facebook’s “like” button, Tumblr’s reblog)
  • These shape what could be expressed

Layer 5: Code and Protocols (Base Layer)

  • HTML/CSS, JavaScript, APIs
  • The technical substrate everything else is built on

Why This Matters:

If you only preserve Layer 1 (content), you lose context. A tweet without timestamp, author, and reply chain is nearly meaningless. A Vine without knowledge that it’s 6 seconds (platform affordance) loses its cultural significance.

Best Practice: Preserve all accessible layers, not just content.

Temporal Stratigraphy: Change Over Time §

Platforms evolve. Preserving multiple snapshots captures this evolution.

Example: Twitter’s Stratigraphy (2006-2025)

Period Character Limit Key Features Cultural Context
2006-2009 140 characters SMS-based, public only Early adopters, tech culture
2010-2013 140 @mentions, hashtags, retweets Mainstream adoption, Arab Spring
2014-2017 140 Embedded images/video, polls Visual turn, meme culture
2017-2022 280 Threads, longer tweets Discourse shift, Trump era
2022-2025 280+ Elon ownership, chaos Decline, exodus to alternatives

If you only archived Twitter in 2025, you’d miss how 140-character limit shaped early Twitter culture. Stratigraphic preservation (periodic snapshots) captures evolution.

Excavating Through Layers: Practical Example §

Scenario: Preserving a Tumblr blog (before or after NSFW purge).

Layer 1 — Content:

  • Use Tumblr’s API or export tool to download posts
  • Format: JSON or HTML
  • Includes text, images, embedded media

Layer 2 — Metadata:

  • Post timestamps (when was this published?)
  • Tags (how did author categorize this?)
  • Post type (text, photo, quote, link, chat, audio, video)
  • Note count (likes + reblogs)

Layer 3 — Relationships:

  • Reblog chain (who reblogged from whom?)
  • Follower graph (if accessible)
  • External links (what other sites/blogs mentioned?)

Layer 4 — Platform Affordances:

  • Tumblr’s “reblog” culture (different from Twitter’s “quote tweet”)
  • Tag system (used for discovery, not just categorization)
  • Dashboard feed (algorithmic? chronological?)

Layer 5 — Code:

  • Tumblr’s HTML theme (custom CSS)
  • Embedded JavaScript (any interactive elements?)

Preservation Strategy:

  1. Download JSON export (Layers 1-2)
  2. Scrape full HTML (captures Layer 4 affordances via design)
  3. Reconstruct reblog chains from metadata (Layer 3)
  4. Document platform context in README (Layer 4 cultural norms)

Part III: Excavation Techniques — Tools and Methods §

Technique 1: Simple Web Scraping (Static Sites) §

Use Case: Static HTML sites (blogs, personal homepages, early web).

Tools:

  • wget: Command-line downloader (recursive crawling)
  • HTTrack: GUI-based website copier
  • ArchiveBox: Modern, full-featured archiver

Example: wget Command

bash
wget --recursive --level=5 --no-parent --wait=1 \
     --convert-links --page-requisites \
     --user-agent="ArchiveBot/1.0" \
     https://example.com

Explanation:

  • --recursive: Follow links
  • --level=5: Crawl up to 5 levels deep
  • --no-parent: Don’t ascend to parent directories
  • --wait=1: Wait 1 second between requests (polite crawling)
  • --convert-links: Make links work offline
  • --page-requisites: Download CSS, images, JavaScript
  • --user-agent: Identify yourself (ethical scraping)

Pros:

  • Fast, simple, reliable for static sites

Cons:

  • Fails on JavaScript-heavy sites (doesn’t execute JS)
  • Can’t handle logins or authenticated content

Technique 2: Browser Automation (Dynamic Sites) §

Use Case: JavaScript-heavy sites (React, Angular apps) or sites requiring interaction.

Tools:

  • Selenium: Browser automation framework
  • Puppeteer: Headless Chrome control (Node.js)
  • Playwright: Modern cross-browser automation

Example: Puppeteer Script (Simplified)

javascript
const puppeteer = require('puppeteer');
const fs = require('fs');

(async () => {
  const browser = await puppeteer.launch();
  const page = await browser.newPage();

  // Navigate to page
  await page.goto('https://example.com/post/12345');

  // Wait for dynamic content to load
  await page.waitForSelector('.post-content');

  // Extract content
  const content = await page.evaluate(() => {
    return {
      title: document.querySelector('.post-title').innerText,
      body: document.querySelector('.post-content').innerText,
      timestamp: document.querySelector('.post-date').innerText
    };
  });

  // Save as JSON
  fs.writeFileSync('post_12345.json', JSON.stringify(content, null, 2));

  await browser.close();
})();

Pros:

  • Handles JavaScript rendering
  • Can simulate user interactions (clicks, scrolls)
  • Can log in to authenticated sites

Cons:

  • Slower than wget (must render pages)
  • More complex to set up

Technique 3: API Harvesting (Structured Data) §

Use Case: Platforms with public APIs (Twitter, Reddit, Mastodon).

Tools:

  • Platform-specific libraries: tweepy (Twitter), PRAW (Reddit)
  • HTTP clients: curl, requests (Python), fetch (JavaScript)

Example: Twitter API (Pre-Elon, when API was good)

python
import tweepy
import json

# Authenticate
auth = tweepy.OAuthHandler(consumer_key, consumer_secret)
auth.set_access_token(access_token, access_token_secret)
api = tweepy.API(auth)

# Download user's timeline
tweets = []
for tweet in tweepy.Cursor(api.user_timeline, screen_name='example_user', tweet_mode='extended').items():
    tweets.append({
        'id': tweet.id_str,
        'text': tweet.full_text,
        'created_at': str(tweet.created_at),
        'retweet_count': tweet.retweet_count,
        'favorite_count': tweet.favorite_count
    })

# Save as JSON
with open('example_user_tweets.json', 'w') as f:
    json.dump(tweets, f, indent=2)

Pros:

  • Structured data (JSON, XML)
  • Includes metadata (timestamps, IDs, relationships)
  • Respects rate limits (built into libraries)

Cons:

  • Platform must have API (many don’t, or shut it down before dying)
  • Rate limits can be restrictive (slow)
  • APIs often sunset before platforms die (Twitter 2023)

Technique 4: Database Extraction (Direct Access) §

Use Case: You have legitimate access to platform’s database (employee, owner, partnership).

Method:

  • SQL dump (if relational database)
  • NoSQL export (if MongoDB, CouchDB, etc.)
  • File system copy (if files stored on disk)

Example: MySQL Dump

bash
mysqldump -u username -p database_name > backup.sql

Pros:

  • Complete, perfect fidelity
  • Includes all metadata, relationships, deleted content

Cons:

  • Rare (requires cooperation from platform)
  • May include sensitive data (must redact)

Technique 5: Forensic Recovery (Post-Mortem) §

Use Case: Platform already shut down, but you might recover fragments.

Methods:

Google Cache:

  • Search cache:example.com in Google
  • Captures recent snapshots (but only for indexed pages)

Wayback Machine:

  • Check Internet Archive’s Wayback Machine
  • May have periodic snapshots

User Backups:

  • Ask former users if they exported their data
  • Crowdsource fragments

Web Archives:

  • Other web archives (UK Web Archive, Library of Congress)

Old Hard Drives:

  • If servers were sold/discarded, forensic data recovery possible (rare, expensive)

Example: MySpace Music Recovery Attempt

After MySpace lost 50 million songs (2019):

  • Internet Archive had some (but not comprehensive)
  • Users who downloaded MP3s shared them
  • Some songs recovered via Google Cache before it expired
  • Most (~90%) permanently lost

Lesson: Forensic recovery is last resort. Success rate is low. Better to preserve proactively.


Part IV: Provenance and Chain of Custody §

Why Provenance Matters §

In physical archaeology, provenance (where an artifact came from) is crucial. An Egyptian vase in a museum is worthless if you don’t know which tomb it came from. Context gives meaning.

In digital archaeology, provenance includes:

  1. Where did you get this? (scraped from live site? downloaded via API? recovered from backup?)
  2. When did you capture it? (date/time of preservation)
  3. Who captured it? (individual, institution, bot)
  4. What modifications were made? (did you redact personal info? convert formats?)

Without provenance, digital artifacts lose credibility.

Chain of Custody Documentation §

Best Practice: Document every step from capture to storage.

Template: Provenance Record

bash
Artifact: GeoCities site "geocities.com/SiliconValley/1234"
Captured: 2009-11-15 03:42 UTC
Method: wget recursive scrape
Captured By: Archive Team volunteer #7823
Source State: Live website (platform still online)
Completeness: 87% (some images 404'd during capture)
Storage: Initial storage on volunteer's hard drive
Transfer: Uploaded to Archive Team torrent 2009-11-20
Current Location: Internet Archive, GeoCities torrent seed
Format: Original HTML + images (no conversion)
Modifications: None (bit-perfect capture)
Verification: MD5 checksums recorded at capture
Access: Public (torrent freely downloadable)

Why Each Field Matters:

  • Captured date: Proves this is snapshot from specific moment
  • Method: Explains why some content might be missing (wget can’t execute JavaScript)
  • Completeness: Honest about gaps (87% is still valuable)
  • Chain of custody: Volunteer → Torrent → Internet Archive (transparent)
  • Modifications: None (proves authenticity)
  • Verification: Checksums prove files unchanged since capture

Metadata Standards §

Use existing standards when possible:

Dublin Core:

  • Standard metadata for digital objects
  • Fields: Creator, Date, Title, Description, Format, Rights

METS (Metadata Encoding and Transmission Standard):

  • Library of Congress standard
  • Used for complex digital objects (multiple files, relationships)

PREMIS (Preservation Metadata):

  • Focuses on provenance and preservation actions
  • Records: who did what, when, why

Example: Dublin Core for Preserved Vine

xml
<metadata>
  <dc:title>Vine #284619204 "Why You Always Lying"</dc:title>
  <dc:creator>Nicholas Fraser (@downgoes.fraser)</dc:creator>
  <dc:date>2015-09-02</dc:date>
  <dc:type>Video (6 seconds, looped)</dc:type>
  <dc:format>MP4 (H.264)</dc:format>
  <dc:description>Viral Vine meme, 30M+ loops, inspired song</dc:description>
  <dc:rights>Fair Use (platform shutdown, cultural preservation)</dc:rights>
  <dc:source>Vine.co (platform shut down 2017)</dc:source>
  <dc:coverage>Internet Archive, Vine Archive Collection</dc:coverage>
  <dc:identifier>IA-Vine-284619204</dc:identifier>
</metadata>

Part V: Ethical Excavation §

Archaeology’s Ethical Evolution §

Physical archaeology has a dark history:

  • Colonial looting (British Museum filled with stolen artifacts)
  • Destroying sites (19th-century excavations were destructive)
  • Ignoring indigenous communities (treating their ancestors as “objects of study”)

Modern archaeology has reformed:

  • Repatriation: Returning artifacts to communities of origin
  • Community consultation: Indigenous peoples have say in excavations
  • Non-destructive methods: Ground-penetrating radar instead of digging
  • Context preservation: Documenting everything, not just taking treasures

Digital archaeology must learn these lessons.

Ethical Principles for Digital Excavation §

1. Minimize Harm

To the platform:

  • Don’t overload servers (respect rate limits)
  • Identify your scraper (user-agent string)
  • Stop if platform asks (honor robots.txt)

Even if you oppose the platform’s business model, harming their infrastructure isn’t ethical (hurts users, not executives).

To users:

  • Don’t expose private information
  • Respect deleted content (if user intentionally deleted, presume they wanted it gone—with exceptions for public figures)

2. Respect robots.txt (Mostly)

robots.txt is a file that tells crawlers what they can/can’t scrape.

Example:

bash
User-agent: *
Disallow: /private/
Disallow: /user-settings/
Allow: /public/

Ethical debate:

  • Strict interpretation: Always obey robots.txt (it’s the site owner’s wishes)
  • Preservation exception: Platform is dying, robots.txt doesn’t apply (saving culture > obeying soon-to-be-dead platform)

Compromise:

  • Respect robots.txt for living platforms
  • Override for dying platforms (with documentation: “We scraped despite robots.txt because platform announced shutdown”)

3. Document Ethical Decisions

When you make ethically contested choices (scraping private content, overriding robots.txt, preserving deleted posts), document why.

Example:

bash
ETHICAL NOTE: Tumblr Post #12345

This post was deleted by the author in 2018 (pre-NSFW purge).
We preserved it because:

1. Author is public figure (political activist with 100k followers)
2. Post documents historically significant event (protest organization)
3. Post was public for 3 years (widely shared, cited in news)

However, we restricted access:

- Not searchable via Google
- Requires researcher credentials to view
- Will honor takedown request if author contacts us

Decision made: 2024-11-15
Decision maker: [Archivist ID]

Transparency builds trust.

4. Community Consultation (When Possible)

If you’re preserving a community’s content (fandom, activist group, cultural community), ask them.

Example: Fan Fiction Archive

Before scraping abandoned LiveJournal fandom:

  1. Post in fandom spaces: “We’re considering archiving [fandom] LiveJournal. Thoughts?”
  2. Listen to concerns (privacy, consent, cultural norms)
  3. Adapt plans (maybe restrict access, or honor individual takedown requests)

Not always possible (no time, or community scattered). But when possible, consultation is ethical.

5. Allow Takedowns

Even after preserving, respect author requests to remove their content.

Process:

  1. Public takedown request form
  2. Verify requester is original author (prevent abuse)
  3. Remove content within reasonable time (7-30 days)
  4. Document removal (provenance: “Content removed 2024-11-20 at author’s request”)

Part VI: Case Study — Excavating a Dying Forum §

Scenario: Small Community Forum (2024) §

Background:

  • Forum: “VintageGamers” (retro gaming community)
  • Software: vBulletin 3.8 (old forum software)
  • Content: 15 years of discussions (2009-2024)
  • Size: 5,000 members, 100,000 posts
  • Announcement: Shutting down in 30 days (hosting costs too high)

Reconnaissance (Days 1-3):

  1. Platform assessment:
  2. vBulletin forum (database-driven)
  3. Public content (no login to read, but login to see images)
  4. URL pattern: /showthread.php?t=[thread_id]
  5. Estimated 10,000 threads

  6. Existing documentation:

  7. Not in Internet Archive (robots.txt blocked crawlers)
  8. No official export tool
  9. Community members panicking, want it saved

  10. Contact admin:

  11. Email forum owner: “Can you provide database dump?”
  12. Owner agrees! (relieved someone cares)
  13. Owner provides MySQL dump (20MB compressed)

Database Excavation (Days 4-7):

  1. Import database locally: bash mysql -u root -p < vintagegamers_backup.sql

  2. Analyze schema:

  3. Tables: posts, threads, users, attachments
  4. Relationships: thread_id links posts to threads
  5. Metadata: timestamps, user IDs, post counts

  6. Export to JSON: ```python import mysql.connector import json

db = mysql.connector.connect(host=”localhost”, user=”root”, password=”…”, database=”vintagegamers”) cursor = db.cursor()

# Export threads cursor.execute(“SELECT thread_id, title, user_id, post_date FROM threads”) threads = [{‘id’: row[0], ‘title’: row[1], ‘user_id’: row[2], ‘date’: str(row[3])} for row in cursor.fetchall()]

with open(‘threads.json’, ‘w’) as f: json.dump(threads, f, indent=2)

# (Repeat for posts, users, etc.) ```

  1. Download attachments:
  2. Images stored in /attachments/ directory
  3. Use wget to download all: bash wget -r -l 1 -A jpg,png,gif https://vintagegamers.com/attachments/

Curation (Days 8-14):

  1. Redact personal info:
  2. Email addresses in user profiles → removed
  3. IP addresses in logs → removed
  4. Private messages → excluded from export

  5. Add metadata:

  6. Create README.md documenting forum history
  7. List notable threads (“Best of VintageGamers”)
  8. Interview longtime members (oral history)

  9. Build search interface:

  10. Use Elasticsearch to index posts
  11. Simple web UI: search by keyword, date, user

Preservation (Days 15-30):

  1. Upload to Internet Archive:
  2. Create “VintageGamers Archive” collection
  3. Upload database dump, JSON exports, attachment images, README

  4. Seed BitTorrent:

  5. Create torrent of full archive
  6. Ensure redundancy (if IA ever goes down)

  7. Announce to community:

  8. Post in forum: “Archive complete! Here’s where to find it.”
  9. Community grateful, downloads personal copies

Provenance Record:

bash
Archive: VintageGamers Forum (2009-2024)
Captured: 2024-11-01 to 2024-11-15
Method: MySQL database dump provided by forum administrator
Captured By: [Archaeobytologist Name], with admin cooperation
Completeness: 100% (full database export)
Redactions: Email addresses, IP addresses, private messages removed
Storage: Internet Archive + BitTorrent
Format: MySQL dump (raw), JSON (parsed), HTML (rendered)
Access: Public (Internet Archive), with restricted personal data
License: CC BY-NC-SA 4.0 (preserves community content, non-commercial)

Outcome:

  • Forum shuts down on schedule
  • 100% of content preserved
  • Community can still access their history
  • Future researchers can study retro gaming community

Conclusion: The Archaeologist’s Mindset §

Digital excavation isn’t just about running scripts. It’s about bringing an archaeological mindset to ephemeral platforms:

1. Systematic: Survey before digging. Plan your excavation. Document everything.

2. Stratigraphic: Preserve all layers (content, metadata, relationships, affordances), not just surface.

3. Contextual: Provenance matters. Where did this come from? When? Who captured it?

4. Ethical: Minimize harm. Respect communities. Be transparent about contested choices.

5. Urgent: Unlike physical archaeology, you don’t have centuries. You have days or weeks. Move fast—but systematically.

In the next chapter, we’ll dive deeper into Digital Forensics—the technical methods for recovering data from damaged, corrupted, or deliberately deleted sources. Sometimes, platforms don’t give you clean MySQL dumps. Sometimes, you’re working with fragments, corrupted files, and deleted evidence.

Digital forensics teaches you how to work with what’s broken.


Discussion Questions §

  1. Methodology: Should digital archaeology prioritize speed (scrape everything quickly) or precision (careful documentation)? How do you balance urgency with rigor?

  2. Ethics: Is it ethical to scrape a platform that explicitly forbids it (robots.txt, ToS) if the platform is dying and content will be lost?

  3. Provenance: Why does documenting where you got an artifact matter? What happens if provenance is unclear or contested?

  4. Stratigraphic Layers: What digital “layers” do you think are most important to preserve? Content? Metadata? Relationships? Platform affordances?

  5. Community Consultation: When is it necessary to consult communities before preserving their content? When is it acceptable to preserve without asking?

  6. Personal Practice: Have you ever “excavated” your own digital artifacts (downloaded Facebook archive, exported tweets)? What did you learn?


Exercise: Plan a Digital Excavation §

Scenario: A platform you use announces shutdown in 60 days. Plan its excavation.

Choose a platform:

  • Small forum you participate in
  • Discord server you’re part of
  • Niche social network
  • Personal blog community

Part 1: Reconnaissance Report (500 words)

  • Platform type, scale, data types
  • Technical barriers (login walls, APIs, rate limits)
  • Existing documentation (Wayback Machine, community wikis)
  • Estimated preservation time and storage

Part 2: Excavation Strategy (800 words)

  • What tools will you use? (wget, Puppeteer, API clients)
  • What layers will you preserve? (content, metadata, relationships)
  • What’s your timeline? (week-by-week plan)
  • How will you handle rate limits or technical barriers?

Part 3: Ethical Framework (500 words)

  • What content should NOT be preserved? (privacy, consent, harm)
  • Will you respect robots.txt? Why/why not?
  • Will you consult the community? How?
  • How will you handle takedown requests?

Part 4: Provenance Documentation (300 words)

  • Write a provenance record for your imagined excavation
  • Include: capture method, date, completeness, modifications, storage

Part 5: Reflection (200 words)

  • What’s the hardest part of this excavation?
  • What ethical dilemmas did you face?
  • Would you actually do this if the platform announced shutdown?

Further Reading §

On Web Archiving Methods §

  • Brügger, Niels. The Archived Web: Doing History in the Digital Age. MIT Press, 2018.
  • How to work with web archives as historical sources

  • Niu, Jinfang. “An Overview of Web Archiving.” D-Lib Magazine 18, no. 3/4 (2012).

  • Technical methods for web preservation

  • Archive Team. “So You Want to Archive a Website.” https://wiki.archiveteam.org/

  • Practical guide from guerrilla archivists

On Digital Forensics §

  • Carrier, Brian. File System Forensic Analysis. Addison-Wesley, 2005.
  • Technical deep dive on recovering deleted data

  • Kirschenbaum, Matthew. Mechanisms: New Media and the Forensic Imagination. MIT Press, 2008.

  • Humanistic approach to digital forensics

On Archaeological Methods (Physical) §

  • Renfrew, Colin, and Paul Bahn. Archaeology: Theories, Methods, and Practice. Thames & Hudson, 2016.
  • Classic textbook (useful for understanding stratigraphic thinking)

  • Hodder, Ian. “The Interpretation of Documents and Material Culture.” In Handbook of Qualitative Research, edited by Norman Denzin and Yvonna Lincoln, 393-402. Sage, 2000.

  • Interpretive archaeology (translates to digital context)

On Ethics §

  • Society for American Archaeology. “Principles of Archaeological Ethics.” https://www.saa.org/
  • Professional ethics code (adaptable to digital)

  • Caswell, Michelle. Urgent Archives: Enacting Liberatory Memory Work. Routledge, 2021.

  • Ethics of archiving marginalized communities

End of Chapter 7

Next: Chapter 8 — Digital Forensics for Archaeobytologists

Part II • Excavation Methods & Digital Forensics

Chapter 8: Digital Forensics for Archaeobytologists

19 min read 4,122 words

Opening: The Crime Scene Is Digital §

In 2019, a hard drive arrived at the Internet Archive. It had been recovered from a dumpster behind a defunct web hosting company. The company had gone bankrupt, its servers sold for scrap, its customer data—thousands of personal websites from the early 2000s—abandoned.

The hard drive was physically intact but logically corrupted. The file system was damaged. File names were mangled or missing. Timestamps were wrong. Some files were partially overwritten with random data. But somewhere in those magnetic sectors were websites that existed nowhere else—personal blogs, family photos, amateur art portfolios. Digital artifacts on the verge of permanent loss.

This required digital forensics: the practice of recovering, analyzing, and authenticating digital evidence from damaged, deleted, or deliberately obscured sources.

Digital forensics emerged from law enforcement (recovering deleted files from criminals’ computers) and IT security (analyzing malware, investigating breaches). But Archaeobytologists need these same skills for different purposes:

  • Recovering deleted content (when users or platforms erase artifacts)
  • Analyzing corrupted files (bit rot, damaged storage media)
  • Authenticating artifacts (proving a file is what it claims to be)
  • Extracting hidden data (metadata, version histories, deleted revisions)
  • Reverse engineering formats (when documentation is lost)

Unlike law enforcement, we’re not building criminal cases. Unlike IT security, we’re not defending against attacks. We’re rescuing cultural artifacts from technological decay.

This chapter teaches digital forensics adapted for Archaeobytological practice. You’ll learn:

  • File system analysis and data recovery
  • Metadata extraction and interpretation
  • Format identification and conversion
  • Authenticity verification and chain of custody
  • Emulation and compatibility layers
  • Ethical boundaries (when forensics becomes invasion)

By the end, you’ll be able to take a corrupted hard drive, deleted website, or mysterious file format and systematically extract whatever cultural value remains.


Part I: Foundations of Digital Forensics §

The Digital Artifact as Evidence §

Physical artifacts are tangible—you can touch a clay pot, examine it with eyes and hands. Digital artifacts are abstract—they’re electromagnetic patterns interpreted by software.

This abstraction creates both challenges and opportunities:

Challenges:

  • Fragility: Flip one bit, and an entire file becomes unreadable
  • Dependency: Files require specific software to interpret (a .doc file is meaningless without Word or a compatible reader)
  • Mutability: Digital files can be silently altered (no visible wear like on physical objects)
  • Ephemerality: Storage media degrades (magnetic fields fade, flash memory loses charge)

Opportunities:

  • Perfect copying: Digital files can be duplicated without loss (unlike physical artifacts)
  • Deep analysis: Can examine file structure bit-by-bit (like x-raying a painting)
  • Metadata richness: Digital files carry embedded information (creation date, author, edit history)
  • Automated processing: Can analyze thousands of files programmatically (impossible with physical artifacts)

The Forensic Workflow §

Digital forensics follows a systematic process:

1. Acquisition (get a copy without altering the original) 2. Preservation (create forensic images, maintain chain of custody) 3. Analysis (examine the data, extract information) 4. Documentation (record findings, methods, provenance) 5. Presentation (make findings accessible to non-technical audiences)

This workflow ensures:

  • Integrity: Original evidence isn’t contaminated
  • Reproducibility: Others can verify your findings
  • Transparency: Methods are documented
  • Legal defensibility: Even though we’re not in court, rigorous methods build credibility

Part II: File System Forensics — Finding the Lost §

Understanding File Systems §

When you delete a file, it doesn’t vanish immediately. The operating system marks the space as “available” but doesn’t erase the data until something overwrites it. This is why “deleted” files can often be recovered.

Common File Systems:

FAT32 (old Windows, USB drives)

  • Simple structure
  • No journaling (prone to corruption)
  • Easy to recover deleted files

NTFS (modern Windows)

  • Complex structure with metadata
  • Journaling (tracks changes, helps recovery)
  • Harder but more sophisticated recovery

ext4 (Linux)

  • Journaling filesystem
  • Can recover recently deleted files from journal

APFS (modern macOS)

  • Encryption by default (complicates recovery)
  • Snapshots (may preserve deleted files)

HFS+ (older macOS)

  • Similar to NTFS in recoverability

Data Carving: Recovering Files Without Metadata §

When file systems are severely damaged (corrupted directory structure, missing file allocation table), you can’t rely on the filesystem to tell you where files are. Instead, you use data carving: scanning raw disk sectors looking for file signatures.

How It Works:

Every file type has a signature (magic bytes) at the beginning:

  • JPEG: FF D8 FF (first three bytes)
  • PNG: 89 50 4E 47 (‰PNG)
  • PDF: 25 50 44 46 (%PDF)
  • ZIP: 50 4B 03 04 (PK..)
  • GIF: 47 49 46 38 (GIF8)

Data carving tools scan the entire disk, looking for these signatures. When found, they extract the file.

Tools:

  • Foremost: Carves files based on headers/footers
  • Scalpel: Fast carving with configurable signatures
  • PhotoRec: Specializes in photos but handles many formats
  • Bulk Extractor: Carves and analyzes (finds emails, URLs, credit cards)

Example: Carving a Corrupted USB Drive

bash
# Install PhotoRec (comes with TestDisk)
sudo apt install testdisk

# Run PhotoRec on drive (replace /dev/sdX with actual device)
sudo photorec /dev/sdX

# Navigate menus:
# 1. Select partition
# 2. Choose file systems to search
# 3. Select output directory
# 4. Wait (can take hours for large drives)

Result: PhotoRec dumps recovered files into output directory, organized by type. Files are renamed generically (f0001.jpg, f0002.png) since metadata is lost.

Limitations:

  • No original filenames (metadata gone)
  • No directory structure (everything dumped together)
  • Fragmented files may be incomplete (if portions were overwritten)
  • Many false positives (random data matching signatures)

Archaeological Application:

When you recover an old hard drive from a defunct web hosting company, data carving may be your only option. You won’t know which files belong to which user or what they were originally named, but you’ll have the actual content—which is better than nothing.

File System Timeline Analysis §

Even when files aren’t deleted, timestamps reveal important information:

MAC Times:

  • Modified: When file content last changed
  • Accessed: When file was last opened
  • Changed: When metadata (permissions, ownership) last changed

Plus NTFS adds:

  • Created: When file was first created

Why Timestamps Matter:

Example 1: Identifying Original Creator

  • A website claims to have been “online since 1998”
  • File timestamps show HTML files created in 2003
  • Either the claim is false, or files were re-uploaded (migration?)
  • Forensic investigation needed

Example 2: Detecting Tampering

  • Archive claims to be “untouched original” from 2005
  • Modified timestamps are 2019
  • Someone edited files after archiving
  • Need to determine what changed

Tools:

  • fls (Sleuth Kit): Lists files with MAC times
  • mactime (Sleuth Kit): Creates timeline from fls output
  • log2timeline/Plaso: Comprehensive timeline analysis

Example: Creating a Timeline

bash
# Install Sleuth Kit
sudo apt install sleuthkit

# Create body file (filesystem metadata)
fls -r -m C: /dev/sdX > bodyfile.txt

# Create timeline
mactime -b bodyfile.txt -d > timeline.csv

# Analyze timeline (Excel, grep, Python)
grep "2009-10" timeline.csv  # Find files from Oct 2009

Archaeological Application:

When analyzing a preserved platform, timeline analysis reveals:

  • When was content created? (chronology of community)
  • When was site last updated? (signs of abandonment)
  • When were files accessed? (usage patterns, popular content)

Part III: Metadata Forensics — The Hidden Stories §

What Is Metadata? §

Metadata is “data about data”—information embedded in files describing their creation, modification, and context.

Types of Metadata:

1. File System Metadata (from OS)

  • Timestamps (created, modified, accessed)
  • Size, location, permissions
  • Captured by filesystem, not embedded in file

2. Embedded Metadata (inside file)

  • EXIF (photos): Camera model, GPS location, date/time
  • ID3 (MP3s): Artist, album, genre, cover art
  • PDF: Author, creation software, edit history
  • Office docs: Author name, organization, edit time, revision history

3. Application Metadata (created by software)

  • HTML: Generator meta tags (<meta name="generator" content="WordPress">)
  • Images: Software used (Photoshop layers, GIMP xcf data)
  • Videos: Codec, bitrate, editing software

Extracting Metadata §

Tool: ExifTool (universal metadata reader)

bash
# Install ExifTool
sudo apt install libimage-exiftool-perl

# Extract all metadata from a file
exiftool photo.jpg

# Extract specific fields
exiftool -CreateDate -Make -Model photo.jpg

# Process entire directory, export to CSV
exiftool -csv -r /path/to/photos/ > metadata.csv

# Remove metadata (privacy scrubbing)
exiftool -all= photo.jpg

Example Output (JPEG from phone):

bash
File Name                       : IMG_2034.jpg
File Size                       : 2.3 MB
File Modification Date/Time     : 2018:11:15 14:23:01
File Type                       : JPEG
EXIF Version                    : 0231
Date/Time Original              : 2018:11:15 14:22:58
Create Date                     : 2018:11:15 14:22:58
Make                            : Apple
Camera Model Name               : iPhone 7
Lens Model                      : iPhone 7 back camera 3.99mm f/1.8
GPS Latitude                    : 37 deg 46' 30.12" N
GPS Longitude                   : 122 deg 25' 9.84" W
GPS Altitude                    : 15 m Above Sea Level

What This Reveals:

  • Photo taken Nov 15, 2018 at 2:22 PM
  • Taken with iPhone 7
  • Location: San Francisco (GPS coordinates)
  • File modified slightly after creation (uploaded? edited?)

Privacy and Metadata §

Ethical Dilemma: Metadata often contains personally identifiable information (PII):

  • GPS coordinates (where someone lives, works, travels)
  • Phone/camera serial numbers (can track individual)
  • Author names, organization names (identity)
  • Full edit history (who touched the file)

Archaeobytologist’s Responsibility:

When preserving:

  • Be aware metadata exists
  • Decide: preserve it (research value) or strip it (privacy)?
  • Document your decision

When publishing:

  • Don’t publish GPS coordinates from personal photos
  • Do preserve GPS for historically significant events (protest locations, disaster sites)
  • Strip metadata from ordinary personal files
  • Keep metadata for public figures, official documents

Case Study: Geotagged Photos from Protests

Photos from 2020 Black Lives Matter protests contain GPS metadata. Should archivists preserve it?

Arguments FOR:

  • Historical record (where protests occurred)
  • Research value (studying protest geography)

Arguments AGAINST:

  • Identifies protesters (could face retaliation)
  • Law enforcement could use for prosecution

Compromise:

  • Preserve photos with GPS
  • Restrict access (research-only, IRB approval)
  • Publish photos with GPS stripped (public version)
  • Aggregate data (publish heatmap of protest locations, not individual coordinates)

Metadata as Provenance §

Metadata helps establish provenance—the history and origin of an artifact.

Example: Authenticating a Leaked Document

Someone claims to have a “leaked internal memo from Facebook, dated 2016.”

Forensic Analysis:

bash
exiftool memo.pdf

Output reveals:

bash
Producer: Microsoft Word 2019
CreateDate: 2021:03:15 09:34:22
ModifyDate: 2021:03:15 09:34:22
Author: John Smith

Findings:

  • Created in 2021 (not 2016)
  • Author listed as “John Smith” (was this Facebook employee? check LinkedIn)
  • Created with Word 2019 (was Word 2019 available in 2016? No—released 2018)

Conclusion: Document is likely fabricated or misdated. Further investigation needed.

Forensic Best Practice:

  • Never trust dates in filenames or document text
  • Check embedded metadata
  • Cross-reference with external evidence (news archives, wayback machine)

Part IV: Format Forensics — Identifying the Unknown §

The Format Identification Problem §

You receive a folder of files from a defunct platform. Many have no file extensions, or wrong extensions (.dat, .tmp, .db). How do you figure out what they are?

Don’t trust extensions. Extensions are metadata (easily changed). Instead, examine the file signature.

Magic Numbers and File Signatures §

Every file format has a magic number—specific bytes at the beginning that identify the type.

Common Signatures:

Format Hex Signature ASCII
JPEG FF D8 FF
PNG 89 50 4E 47 0D 0A 1A 0A ‰PNG....
GIF 47 49 46 38 GIF8
PDF 25 50 44 46 %PDF
ZIP 50 4B 03 04 PK..
MP3 49 44 33 or FF FB ID3 or ÿû
EXE 4D 5A MZ
SQLite 53 51 4C 69 74 65 20 66 6F 72 6D 61 74 20 33 00 SQLite format 3.

Tool: file command (Unix)

bash
# Identify file type
file unknown_file.dat
# Output: unknown_file.dat: PNG image data, 800 x 600, 8-bit/color RGB, non-interlaced

# Check multiple files
file *

Tool: DROID (UK National Archives)

  • GUI tool for format identification
  • Uses PRONOM registry (comprehensive format database)
  • Generates reports on entire directories

Obsolete and Proprietary Formats §

The Hardest Cases:

1. Proprietary formats with no documentation

  • Company went bankrupt, format specs lost
  • Example: Lotus 123 spreadsheets (.wk1, .wk3)

2. Custom binary formats

  • Platform created its own format for efficiency
  • Example: Vine’s proprietary video container

3. Encrypted or obfuscated formats

  • DRM-protected files
  • Example: iTunes FairPlay (before DRM removal)

Strategies:

A. Search for Format Documentation

  • Archive.org (old software manuals)
  • FileFormat.info
  • “Just Solve the File Format Problem” wiki
  • Ask old forums, mailing lists

B. Reverse Engineer

  • Hex editor: examine file structure
  • Strings command: extract readable text
  • Binwalk: analyze binary structure
  • Compare multiple examples to find patterns

C. Find Old Software

  • Run original software in emulator
  • Example: Run MS-DOS programs in DOSBox to open ancient file formats

D. Convert via Emulation

  • Open file in original software, export to modern format
  • Lossy but better than nothing

Example: Recovering WordPerfect 5.1 Documents

WordPerfect was dominant in 1980s-90s. Many legal documents, dissertations, novels exist only in .wpd format.

Solution:

  1. Download WordPerfect 5.1 (abandonware)
  2. Run in DOSBox emulator
  3. Open .wpd files
  4. Export to ASCII or RTF (WordPerfect can do this)
  5. Import to modern word processor

Alternative: LibreOffice can open some WordPerfect formats (but not perfectly).


Part V: Emulation and Compatibility §

When Files Require Specific Environments §

Some digital artifacts aren’t just files—they’re experiences that require specific software, hardware, or operating systems.

Categories:

1. Software Applications

  • Need specific OS (Windows 95 programs won’t run on modern Windows)
  • Example: Old educational CD-ROMs

2. Websites with Complex JavaScript

  • Need specific browser versions
  • Example: Flash-based sites (need Flash Player)

3. Games

  • Need specific hardware (arcade machines, consoles)
  • Example: 1980s arcade games on custom boards

4. Interactive Art

  • Need specific plugins, environments
  • Example: Java applets, Shockwave

Emulation Strategies §

Strategy 1: OS Emulation

Run the entire original operating system in a virtual machine.

Tools:

  • VirtualBox: Run Windows XP, Linux, older systems
  • QEMU: Low-level emulation, supports many architectures
  • DOSBox: Emulate MS-DOS (for 1980s-90s software)

Example: Running Windows 95 Software

  1. Download Windows 95 ISO (abandonware/legally gray)
  2. Create VirtualBox VM
  3. Install Windows 95
  4. Install old software (from CD image or floppy disk image)
  5. Take VM snapshot (preserve working state)
  6. Users can run VM, experience software as originally intended

Strategy 2: Browser-Based Emulation

Internet Archive’s approach: run emulators in web browser.

Technologies:

  • Emularity: JavaScript-based emulation framework
  • JSMESS: Arcade/console emulator in JavaScript
  • Ruffle: Flash emulator in WebAssembly

Example: Internet Archive’s Software Collection

  • Visit: archive.org/details/softwarelibrary
  • Click any old program
  • Emulator loads in browser
  • Run 1980s software without installing anything

Strategy 3: Format Migration

Convert old formats to modern equivalents (lossy but pragmatic).

Examples:

  • Flash → HTML5 (recreate interactions in modern web tech)
  • QuickTime → MP4 (convert video codec)
  • WordPerfect → DOCX (lose some formatting but preserve text)

Trade-offs:

  • Emulation: High fidelity, but requires maintaining emulators
  • Migration: Lower fidelity, but content accessible in modern tools

Best Practice: Do both when possible. Preserve original + create migrated version.


Part VI: Authentication and Chain of Custody §

Proving a Digital Artifact Is Authentic §

Physical artifacts can be authenticated through material analysis (carbon dating, paint chemistry). Digital artifacts are perfectly copyable—a copy is identical to original. So how do you prove authenticity?

Cryptographic Hashing §

A hash is a unique fingerprint of a file. Change one bit, and the hash changes completely.

Common Hash Functions:

  • MD5: 128-bit hash (fast but cryptographically broken—don’t use for security)
  • SHA-1: 160-bit hash (deprecated, collisions found)
  • SHA-256: 256-bit hash (current standard)
  • SHA-512: 512-bit hash (even stronger)

Example: Computing SHA-256 Hash

bash
# Hash a single file
sha256sum file.jpg
# Output: a1b2c3d4e5f6... file.jpg

# Hash all files in directory
find . -type f -exec sha256sum {} \; > manifest.txt

# Verify files haven't changed
sha256sum -c manifest.txt
# Output: file.jpg: OK

Use Cases:

1. Proving Integrity

  • Archive Team publishes GeoCities torrent with SHA-256 hashes
  • You download torrent, compute hashes
  • If they match, you know data wasn’t corrupted in transit

2. Detecting Tampering

  • Hash preserved website when first captured
  • Years later, re-hash to verify nothing changed
  • If hash differs, investigate (bit rot? deliberate alteration?)

3. Chain of Custody

  • Hash original source
  • Hash after each processing step (conversion, migration)
  • Document all hashes
  • Proves artifact’s history

Digital Signatures §

For legally significant documents, cryptographic signatures prove:

  • Who created/signed the document
  • When it was signed
  • That it hasn’t been altered since signing

How It Works:

  1. Author creates document
  2. Author signs with private key (generates signature)
  3. Anyone can verify signature with author’s public key
  4. Signature proves: (a) author had private key, (b) document unchanged

Tools:

  • GnuPG: Sign and verify documents
  • OpenSSL: Cryptographic operations
  • Adobe Acrobat: PDF signatures

Archaeological Application:

When archiving controversial or historically important documents (leaked memos, government records, deleted tweets), sign them immediately. This proves:

  • You had the document at time of signing
  • Document hasn’t been altered since
  • Protects against accusations of fabrication

Part VII: Forensic Documentation §

Recording Your Process §

Forensic work is worthless if you can’t explain what you did. Document everything:

Acquisition Documentation §

Record:

  • Source device (hard drive model, serial number)
  • Date/time acquired
  • Who acquired it (chain of custody)
  • Tools used (software versions)
  • Hashes (original source)

Example Log:

bash
=== Forensic Acquisition Log ===
Date: 2024-11-15
Examiner: Jane Smith
Case: GeoCities Hard Drive Recovery

Source Device:
  Make: Western Digital
  Model: WD5000AAKS
  Serial: WD-XXXX1234
  Capacity: 500GB

Acquisition Method:
  Tool: dd (GNU coreutils 8.32)
  Command: dd if=/dev/sdb of=geocities_hdd.img bs=4M status=progress
  Duration: 3 hours 42 minutes

Verification:
  SHA-256 (source): a1b2c3d4...
  SHA-256 (image):  a1b2c3d4...
  Match: YES

Notes:

  - Drive had bad sectors (dd_rescue used to skip)
  - Approximately 2.3% of drive unreadable
  - Bad sectors logged in bad_sectors.txt

Analysis Documentation §

Record:

  • What you found
  • How you found it (specific commands, tools)
  • Screenshots (visual proof)
  • Interpretation (what does this mean?)

Example Analysis Notes:

bash
File: mystery_file.dat
Location: /recovered_data/sector_2314/mystery_file.dat

1. Format Identification
   Command: file mystery_file.dat
   Result: "SQLite 3.x database"

2. Schema Analysis
   Command: sqlite3 mystery_file.dat ".schema"
   Result: Tables: users, posts, comments

3. Content Extraction
   Command: sqlite3 mystery_file.dat "SELECT * FROM posts LIMIT 10"
   Result: 10 rows exported to sample.csv

4. Interpretation
   This appears to be a forum database. Contains:

   - 12,342 users
   - 45,678 posts
   - 123,456 comments
   Dates range from 2004-03-15 to 2009-08-22

5. Conclusion
   Likely a phpBB or vBulletin forum database.
   Requires further analysis to identify specific platform.

Part VIII: Ethical Boundaries in Forensics §

When Forensics Becomes Invasion §

Digital forensics is powerful—but power requires ethical limits.

Scenarios Where Forensics Is Inappropriate §

1. Private Communications

  • Deleted emails, DMs, chats
  • Just because you can recover them doesn’t mean you should

2. Intimate Content

  • Personal photos, videos, journals
  • Respect people’s decision to delete

3. Trade Secrets / Proprietary Information

  • Corporate data on abandoned servers
  • May be legally protected even if physically accessible

4. Ongoing Harm

  • Harassment campaigns, doxxing, revenge porn
  • Forensic recovery could perpetuate harm

Forensic Ethics Framework §

Ask before analyzing:

1. Consent

  • Did creator consent to preservation?
  • Can you obtain consent now?

2. Public Interest

  • Is this historically/culturally significant?
  • Does public value outweigh privacy concerns?

3. Harm Potential

  • Could forensic recovery cause harm?
  • To whom? How severe?

4. Alternative Methods

  • Can you achieve your goal without forensics?
  • Is less invasive method available?

Example: The Deleted Political Tweet

A politician deletes a tweet. You have forensic tools to recover it from cached data. Should you?

Analysis:

  • Public figure: Yes (higher scrutiny justified)
  • Public interest: If tweet is newsworthy, yes
  • Harm: Minimal (politician chose public platform)
  • Alternatives: Check Wayback Machine, Politwoops (already doing this)

Conclusion: Ethical to recover and publish (accountability > privacy for public officials).

Example: The Abandoned Teenager’s Blog

You recover a hard drive with a teenager’s private blog from 2005 (they’re now 35). Should you publish it?

Analysis:

  • Private person: Higher privacy expectation
  • Consent: Can’t easily contact them
  • Public interest: Low (unless exceptional historical value)
  • Harm: Could embarrass them (teenage writing)

Conclusion: Don’t publish without consent. Document that it existed, archive privately, contact them if possible.


Conclusion: The Forensic Archaeobytologist §

Digital forensics transforms you from passive archivist to active investigator. You don’t just accept what platforms give you—you dig deeper, recover what was lost, authenticate what’s dubious, and extract meaning from the opaque.

Every corrupted hard drive, every deleted file, every mysterious format is a puzzle. Your forensic skills determine whether that puzzle is solved or remains forever mysterious.

But with great power comes great responsibility. Forensics can invade privacy, resurrect deliberately forgotten content, and cause harm. The Custodial Filter applies here too: just because you can recover something doesn’t mean you should.

In the next chapter, we’ll explore the ethics of preservation in depth—examining the hardest dilemmas Archaeobytologists face, and building frameworks for navigating them.

For now, practice your forensic skills. Find an old hard drive, a corrupted file, a mysterious binary. Apply these methods. Document your process. And ask: What stories are hidden in these bits?

The artifacts are waiting. Now go uncover them.


Discussion Questions §

  1. Metadata Privacy: You’re archiving a photo collection from a defunct platform. GPS coordinates reveal protesters’ locations. Do you strip the metadata or preserve it for research?

  2. Format Obsolescence: You find files in a proprietary format with no documentation. Do you spend weeks reverse-engineering it, or accept that some content will be lost?

  3. Deleted Content: A user intentionally deleted their account and content. You have a backup. Do you preserve it?

  4. Authentication: Someone claims a document is a “leaked corporate memo.” Your forensic analysis shows metadata inconsistencies. How do you publish your findings without enabling misinformation?

  5. Emulation vs. Migration: Is it better to maintain perfect fidelity through emulation (expensive, complex) or accept some loss through format migration (pragmatic, sustainable)?

  6. Chain of Custody: How do you prove to skeptics that an archived artifact is authentic and unaltered?


Exercise: Forensic Recovery Project §

Task: Conduct a forensic analysis of a digital artifact.

Part 1: Acquire an Artifact (Choose one)

  • Old USB drive from a drawer
  • Downloaded corrupt file from internet
  • Deleted file from your own computer (practice recovery)
  • Mystery file with wrong/missing extension

Part 2: Forensic Analysis (1000 words)

Document:

  1. Acquisition: How did you obtain it? Document device info, date, method
  2. Hashing: Compute SHA-256, document hash
  3. Format Identification: What type of file? Use file command or DROID
  4. Metadata Extraction: What metadata exists? Use ExifTool
  5. Content Analysis: What’s inside? Can you open it? Recover data?
  6. Timeline: When was it created, modified, accessed?
  7. Findings: What did you learn? Any surprises?

Part 3: Ethical Reflection (500 words)

  • Was this analysis ethical?
  • Did you encounter private information?
  • How would you handle this if archiving for public access?
  • What would you do differently?

Part 4: Documentation (Create forensic report)

  • Professional-style report documenting your process
  • Include: acquisition log, tool commands, screenshots, findings, conclusions

Further Reading §

On Digital Forensics Methods §

  • Carrier, Brian. File System Forensic Analysis. Addison-Wesley, 2005.
  • Comprehensive technical reference on filesystem analysis

  • Casey, Eoghan. Digital Evidence and Computer Crime. Academic Press, 2011.

  • Forensic investigation methodology

  • Jones, Keith, et al. Real Digital Forensics: Computer Security and Incident Response. Addison-Wesley, 2005.

  • Practical forensics for investigators

On Format Preservation §

  • Brown, Adrian. Practical Digital Preservation. Facet Publishing, 2013.
  • Format identification, migration, emulation strategies

  • Kirschenbaum, Matthew. Mechanisms: New Media and the Forensic Imagination. MIT Press, 2008.

  • Theoretical foundation for digital forensics in humanities

On Emulation §

  • Rosenthal, David S. H. “Emulation & Virtualization as Preservation Strategies.” Report for Mellon Foundation, 2015.
  • Technical and institutional challenges of emulation

  • Internet Archive. “Software Preservation.” https://archive.org/details/softwarelibrary

  • Practical examples of browser-based emulation

Tools Documentation §

  • The Sleuth Kit: http://www.sleuthkit.org/
  • ExifTool: https://exiftool.org/
  • Autopsy (GUI for Sleuth Kit): https://www.autopsy.com/
  • DROID: https://digital-preservation.github.io/droid/

End of Chapter 8

Next: Chapter 9 — The Custodial Filter: Ethics of Preservation

Part II • Excavation Methods & Digital Forensics

Chapter 9: The Custodial Filter

Ethics of Preservation

24 min read 5,089 words

Opening: The Archivist’s Dilemma §

In 2018, a digital archivist received an anonymous hard drive in the mail. No return address, no note. Just a drive containing 50GB of data from a defunct online forum for survivors of domestic abuse.

The forum had shut down three years earlier when its volunteer admin burned out. No backup was ever released publicly. The community scattered, their stories lost. Until this drive appeared.

The archivist faced impossible questions:

Should she preserve this?

  • For: This is crucial documentation of survivor experiences, mutual aid networks, and trauma recovery
  • Against: People shared deeply personal stories under usernames, expecting privacy and eventual deletion

If she preserves it, who gets access?

  • Open access: Researchers, journalists, the public—but risks outing survivors, exposing vulnerabilities
  • Restricted access: Researchers only—but who decides who qualifies?
  • No access: Preserve but seal for 50 years—but then why preserve at all?

Can she even contact the original posters?

  • Most used pseudonyms
  • Forum email addresses are dead
  • No way to get consent

What about the abusers mentioned in posts?

  • Some are named explicitly
  • Preserving could be evidence—or could be defamation
  • Do alleged abusers have privacy rights?

She sat with this drive for months, paralyzed. Every choice felt wrong.

This is the custodial burden—the weight of deciding what gets remembered and what gets forgotten, who gets privacy and who gets accountability, what serves history and what causes harm.

This chapter provides frameworks for navigating these impossible choices. Not easy answers (there are none), but ethical methodologies for thinking through preservation dilemmas systematically.


Part I: The Philosophy of Custodianship §

What It Means to Be a Custodian §

When you preserve a digital artifact, you become its custodian—responsible for its care, its interpretation, and its future.

This isn’t neutral work. Preservation is an ethical act that:

1. Shapes Memory

  • What you save becomes the historical record
  • What you don’t save is forgotten
  • You’re deciding what future generations can know

2. Distributes Power

  • Preservation gives voice to some, denies it to others
  • Archives have historically elevated powerful voices, erased marginalized ones
  • Your choices can reproduce or resist these patterns

3. Affects Living People

  • Digital artifacts often involve people still alive
  • Preservation can help (documentation, accountability) or harm (privacy violations, retraumatization)
  • You must weigh these consequences

4. Encodes Values

  • Every triage decision reflects what you think matters
  • Cultural significance, consent, public interest—these are value judgments
  • Your ethics become embedded in the archive

The Custodial Paradox §

Custodianship involves fundamental tensions:

Preservation vs. Privacy

  • Historians want everything preserved; individuals want to be forgotten
  • Both have legitimate claims

Accountability vs. Compassion

  • Preserving evidence holds wrongdoers accountable
  • But people change; permanent records can be punitive

Comprehensiveness vs. Harm Reduction

  • Saving everything maximizes historical value
  • But some artifacts cause ongoing harm by existing

Present Consent vs. Future Value

  • People may not want something preserved now
  • But it might be historically crucial in 50 years

There are no formulas for resolving these tensions. Only frameworks for deliberation.


Part II: The Custodial Filter — Five Ethical Questions §

The Custodial Filter (introduced in Chapter 5) is a systematic approach to preservation ethics. Before preserving any artifact, ask:

Question 1: Cultural Significance §

Does this artifact matter?

This seems simple but is deeply contested. Who decides what “matters”?

The Traditional Canon Problem

Historically, archives preserved:

  • Elite voices (wealthy, educated, politically powerful)
  • Official records (government, corporations)
  • Dominant cultures (Western, white, male perspectives)

Marginalized communities were systematically erased:

  • LGBTQ+ lives (destroyed as “obscene”)
  • Indigenous knowledge (dismissed as “primitive”)
  • Working-class culture (seen as “low value”)
  • Women’s private writings (deemed “trivial”)

Result: History is biased toward the powerful. Archives reflect and reinforce this.

Decolonizing Significance

New approach: Prioritize voices historically excluded:

  • Marginalized communities (LGBTQ+, disabled, immigrant, Indigenous)
  • Grassroots movements (mutual aid, activism, subcultures)
  • Everyday life (not just “important” people)
  • Dissent (voices challenging power)

Principle: If an artifact documents an underrepresented community or challenges dominant narratives, significance increases.

Community-Defined Significance

Best practice: Ask the community that created content whether it matters.

Example: Trans Archive Project (Hypothetical)

An archivist wants to preserve trans people’s early YouTube videos (2006-2010). Before doing so, she:

  1. Contacts trans creators (where possible)
  2. Asks trans community members whether this is valuable
  3. Prioritizes what the community identifies as significant (not what she assumes)

Result: Preserves what trans people themselves think matters, not what outsiders imagine.

Challenge: Communities aren’t monolithic. Trans people will disagree about what’s important. Archivist must navigate plural perspectives.

Significance Over Time

What seems trivial today may be crucial tomorrow:

  • Social media posts seem ephemeral, but document social movements (Arab Spring, #BlackLivesMatter)
  • Memes seem silly, but reflect cultural anxieties and political discourse
  • Personal blogs seem niche, but record lived experiences

Principle: When uncertain, over-preserve. You can always restrict access later, but you can’t un-lose deleted data.

Did the creator agree to preservation?

This is the hardest question because digital culture blurs public/private boundaries.

Explicit Consent:

  • Creator explicitly licensed work for reuse (Creative Commons, public domain)
  • Creator posted on platform with preservation-friendly TOS
  • Creator contacted and agreed to archiving

Implied Consent:

  • Content posted publicly on open web
  • Platform TOS mentioned archiving (even if users didn’t read it)
  • Content was public for years before platform died

Ambiguous:

  • Content was public but creator expected ephemerality (tweets, Snapchat stories)
  • Content was “friends-only” but platform made it hard to truly restrict access
  • Creator deleted it, but copies survived elsewhere

No Consent / Violated Consent:

  • Private content leaked without permission
  • Content creator explicitly deleted (signal of wanting it forgotten)
  • Content posted in contexts with strong privacy norms (support groups, medical forums)

Sometimes, preserving without consent is justified:

1. Public Figures and Accountability

Politicians, CEOs, and public figures have reduced privacy expectations:

  • Their statements are newsworthy
  • Public has interest in holding them accountable
  • Deleting tweets shouldn’t erase the record

Example: Politician tweets racist statement, then deletes it. Archiving without consent is justified—accountability trumps desire to forget.

2. Historical Significance

Sometimes historical value outweighs individual privacy:

  • Documentation of major events (9/11, Arab Spring)
  • Evidence of corporate or government wrongdoing
  • Records of marginalized communities (with care)

Guideline: The more significant, the more consent can be overridden—but never lightly.

3. Abandoned Content

If creator is unreachable (platform dead, email bounces, user vanished):

  • Reasonable assumption: content is effectively abandoned
  • Preservation prevents loss
  • But: Add opt-out mechanism (if someone emerges claiming it, they can request removal)

1. Private Content Leaked

Hacked emails, leaked DMs, stolen nudes—these should not be preserved, even if newsworthy:

  • Privacy violation is harm
  • Preserving perpetuates harm
  • Journalistic value doesn’t justify violation

Exception: If content reveals serious wrongdoing (corruption, abuse) and no other evidence exists—then restricted-access preservation with redactions might be justified. Case-by-case.

2. Content Explicitly Deleted

If a creator intentionally deleted something (not platform-deleted), that’s a signal:

  • They regret it
  • They want it forgotten
  • Preserving against their will is disrespectful

Exception: Public figures, accountability cases (as above)

3. Vulnerable Populations

Children, abuse survivors, people in crisis—their consent is especially important:

  • Power imbalances may have coerced original posting
  • Ongoing harm from exposure (stalking, harassment)
  • Trauma from being unable to escape past

Guideline: Err on the side of respecting deletion/privacy for vulnerable people.

Ideal: Get explicit consent from everyone. In practice, this is often impossible:

  • Platforms shut down quickly (no time to contact thousands of users)
  • Users are pseudonymous (can’t find them)
  • Users are dead

Pragmatic approach:

  1. Prioritize consent when possible (contact creators if reachable)
  2. Assume implied consent for truly public content (but allow opt-out)
  3. Restrict access for ambiguous cases (preserve but don’t make public)
  4. Don’t preserve clear violations (leaked private content, explicit deletion by vulnerable people)

Question 3: Harm Assessment §

Does preserving this artifact cause harm?

Some content, even if historically significant, causes ongoing harm by continuing to exist.

Types of Harm

1. Direct Physical Harm

Content that enables violence:

  • Doxxing (addresses, phone numbers enabling stalking)
  • Revenge porn (non-consensual intimate images)
  • Terrorist manifestos with actionable plans
  • Harassment campaigns coordinating attacks

Principle: Do not preserve content that directly facilitates physical harm, even if “historically significant.”

2. Psychological Harm

Content that retraumatizes:

  • Graphic violence (mass shooting videos, lynchings)
  • Child sexual abuse material (never preserve, illegal)
  • Intimate details of trauma shared in private contexts, now exposed

Guideline: Preserve metadata (that it existed, summary of what it was) but not the content itself. Document without reproducing harm.

3. Reputational Harm

Old content that unfairly damages someone:

  • Youthful mistakes preserved forever (teenagers doing dumb things)
  • False accusations or rumors
  • Outdated views the person has disavowed

Trade-off: Reputational harm vs. accountability

  • If person is public figure and content shows pattern of behavior → preserve (accountability)
  • If person is private individual and content is isolated incident → consider deletion (compassion)

4. Systemic Harm

Content that perpetuates oppression:

  • Hate speech that normalizes violence against marginalized groups
  • Misinformation that undermines public health (anti-vax, COVID denial)
  • Propaganda that radicalizes (extremist recruitment material)

Complex Calculation:

  • Preserving for research (understanding radicalization) has value
  • But making it accessible can spread harm
  • Solution: Very restricted access, redactions, content warnings

Harm Mitigation Strategies

If you decide to preserve harmful content (for historical/research value), mitigate harm:

1. Restricted Access

  • Researchers only (require IRB approval)
  • Time embargo (seal for X years)
  • Gated access (application process, vetting)

2. Redactions

  • Remove personal information (addresses, phone numbers)
  • Blur faces in videos/photos
  • Anonymize user names (if not public figures)

3. Contextualization

  • Content warnings (trigger warnings for traumatic material)
  • Historical context (explain why this existed, what it reveals)
  • Counter-narratives (provide resources challenging harmful content)

4. Opt-Out Systems

  • Allow people to request removal
  • Regularly review and honor takedown requests
  • Transparent process for appeals

Example: Hate Forum Archive

A researcher wants to preserve a white supremacist forum (to study radicalization):

Harm Assessment:

  • Direct harm: Forum coordinated harassment (yes, harmful)
  • Psychological harm: Racist content traumatizes targets (yes)
  • Systemic harm: Normalizes white supremacy (yes)

Preservation Decision: Yes, but highly restricted

  • Archive content (research value)
  • Redact personal info of victims
  • Researcher access only (IRB required)
  • Content warnings throughout
  • Provide to hate-monitoring orgs (ADL, SPLC) but not public

Harm Mitigation:

  • Not searchable by Google (no SEO amplification)
  • Context: explain why preserved, what it reveals about extremism
  • Counter-resources: link to deradicalization materials

Question 4: Redundancy §

Is someone else already preserving this?

If multiple institutions have copies, your effort might be better spent elsewhere.

Checking for Redundancy

Internet Archive’s Wayback Machine:

  • Search for URL: has it been crawled?
  • How many snapshots? How recent?
  • Are snapshots complete (images, JavaScript, embedded media)?

Library of Congress Web Archive:

  • US government sites, some social media (Twitter)
  • Check their collections

University Archives:

  • Many universities archive specific topics (LGBTQ+ history, political movements)
  • Contact university libraries

Community Archives:

  • Fan archives, activist archives, diaspora archives
  • Often underfunded but comprehensive within niche

Individual Creators:

  • Did creators back up their own content?
  • Many YouTubers, bloggers, podcasters have local copies

When Redundancy Is Valuable

Even if something is archived elsewhere, you might still preserve if:

1. Different Preservation Methods

  • Internet Archive: breadth (millions of sites, shallow scraping)
  • You: depth (one community, rich metadata, contextualization)

2. Institutional Fragility

  • If existing archive is at risk (unstable organization, no long-term funding)
  • Redundancy = resilience (LOCKSS: “Lots of Copies Keep Stuff Safe”)

3. Access Differences

  • Existing archive is restricted; yours could be open
  • Or vice versa: existing archive is too open; yours provides privacy protections

4. Format/Quality Differences

  • Existing archive has low-quality captures (missing images, broken interactivity)
  • You can improve preservation fidelity

When to Defer

If artifact is already well-preserved by stable institutions with good access:

  • Deprioritize (spend your time on endangered, unarchived material)
  • Contribute to existing effort (add metadata, fix errors) rather than duplicate

Principle: Maximize coverage of endangered material; minimize duplication of secure material.

Question 5: Feasibility and Resource Allocation §

Can you realistically preserve this, and is it the best use of resources?

Ethics isn’t just about right vs. wrong—it’s about triage under scarcity.

Resource Constraints

You have limited:

  • Time (especially in crisis preservation)
  • Storage (servers, hard drives cost money)
  • Expertise (technical skills for complex preservation)
  • Legal capacity (some preservation risks lawsuits)

Ethical question: Given these limits, how do you allocate effort?

The Trolley Problem of Triage

Scenario: You have 48 hours before a platform shuts down. You can:

Option A: Preserve 10,000 posts from marginalized creators (high cultural value, small volume)

Option B: Preserve 1,000,000 posts representative sample of entire platform (lower per-post value, but comprehensive dataset)

Option C: Preserve 100 at-risk posts (doxxing targets, abuse survivors) that will cause harm if platform dies and they lose control of deletion

Which do you choose?

No right answer. Depends on:

  • Your mission (community-focused? Comprehensive? Harm reduction?)
  • Others’ efforts (is anyone doing A, B, or C?)
  • Your skills (do you have tools for bulk scraping? Or deep curation?)

Ethical Triage Principles

1. Prioritize the Endangered

  • Things no one else is saving > things already archived
  • Imminently disappearing > stable but declining

2. Prioritize the Unrepresented

  • Marginalized voices > mainstream voices (mainstream is already over-preserved)

3. Prioritize Harm Reduction

  • If failure to preserve causes direct harm (loss of evidence, destruction of community records), prioritize

4. Be Transparent About Trade-offs

  • Document what you chose NOT to save and why
  • Let others second-guess your decisions (transparency enables correction)

5. Accept Imperfection

  • You will make mistakes
  • Some precious artifacts will be lost
  • This is tragic but unavoidable

Grief is part of the work.


Part III: Case Studies in Custodial Ethics §

Case Study 1: The Tumblr NSFW Purge (2018) §

Background:

  • Tumblr banned all “adult content” (2018)
  • Millions of posts deleted (LGBTQ+ content, art, sex education, sex work portfolios)
  • 48-hour warning before purge began

Custodial Dilemma:

Should archivists preserve purged content?

Arguments FOR:

  • Cultural significance (LGBTQ+ history, sex-positive community)
  • Censorship resistance (corporate shouldn’t decide what’s “obscene”)
  • Creators losing work (artists, educators, sex workers losing portfolios)

Arguments AGAINST:

  • Consent ambiguous (some creators chose not to self-archive, signal they wanted it gone?)
  • Adult content has complex consent (performers may not want redistribution)
  • Legal risk (some purged content may have been illegal, archivists don’t want liability)

What Actually Happened:

  • Some archivists saved portions (restricted access, research use only)
  • Many creators self-archived (exported own blogs before purge)
  • Much was permanently lost

Ethical Assessment:

  • Mistake: Archivists were too cautious (legal fears prevented rescue)
  • Lesson: Should have preserved more aggressively, with restricted access
  • Better approach: Preserve everything, then tier access (public for general posts, restricted for adult content, opt-out for anyone requesting)

Case Study 2: The January 6 Insurrection Videos (2021) §

Background:

  • Capitol insurrection (Jan 6, 2021)
  • Participants livestreamed and posted videos
  • Many later deleted content (realizing it was incriminating evidence)

Custodial Dilemma:

Should archivists preserve deleted insurrection videos?

Arguments FOR:

  • Historical significance (major political event)
  • Accountability (participants committed crimes, videos are evidence)
  • Public interest (understanding extremism, documenting attempted coup)

Arguments AGAINST:

  • Creators deleted them (wanted them forgotten)
  • Privacy (even criminals have some privacy rights?)
  • Amplification (preserving could glorify insurrection)

What Actually Happened:

  • ProPublica, FBI, and others archived extensively
  • Parler (platform used) was scraped before it went offline
  • Videos used as evidence in prosecutions

Ethical Assessment:

  • Correct: Public figures committing crimes have no privacy expectation
  • Accountability trumps deletion (deleted evidence doesn’t erase wrongdoing)
  • Historical value high (future generations must understand this event)

BUT: Nuance Required

  • Bystanders’ faces should be blurred (not all participants, some just present)
  • Victims’ identities protected (Capitol police officers, staff)
  • Context added (this was a coup attempt, not a legitimate protest)

Case Study 3: The GeoCities Rescue (2009) §

Background:

  • GeoCities shutdown (2009, 3 weeks warning)
  • Archive Team scraped 650GB (fraction of total)
  • No time to contact creators for consent

Custodial Dilemma:

Preserving without consent—ethical?

Arguments FOR:

  • Cultural significance (early web history, millions of voices)
  • Abandonment (most sites hadn’t been updated in years; creators gone)
  • Historical value (documenting 1990s-2000s internet culture)

Arguments AGAINST:

  • No consent (couldn’t contact millions of users)
  • Privacy violations (personal info, old photos, embarrassing content)
  • Context collapse (sites made for small audiences, now exposed to anyone)

What Actually Happened:

  • Archive Team scraped publicly accessible sites
  • Posted as downloadable torrent
  • Many former GeoCities users grateful (recovered lost memories)
  • Some users upset (wanted sites to die with platform)

Ethical Assessment:

  • On balance, correct: Historical value high, consent impossible to obtain, default to preservation
  • Could improve: Better opt-out system (allow people to request removal from torrent/online archives)
  • Lesson: When consent is impossible and significance is high, preserve—but build in takedown processes

Case Study 4: The Survivor Forum Hard Drive (Opening Scenario) §

Background: Domestic abuse survivor forum (defunct 3 years), hard drive anonymously mailed to archivist

Custodial Dilemma:

What should archivist do?

Options:

Option A: Destroy

  • Private content, no consent, trauma risk
  • Survivors have right to privacy
  • Let it die with the forum

Option B: Preserve but Seal

  • Lock it away for 50 years
  • Protects privacy now, allows future access
  • But: Why preserve if no one can use it?

Option C: Restricted Research Access

  • Give to trauma researchers, domestic violence orgs
  • Could help others, inform policy
  • But: Still violates privacy of posters

Option D: Try to Contact Posters

  • Track down users (if possible), ask consent
  • Respect their wishes individually
  • But: Contacting could retraumatize, or alert abusers to their mentions

Ethical Analysis:

Competing Values:

  • Privacy (posters expected confidentiality)
  • Historical value (documents survivor experiences, mutual aid networks)
  • Potential benefit (research could help other survivors)
  • Harm risk (exposure could endanger people)

Recommendation: Option B or C (Preserve but Restrict)

Reasoning:

  1. Destroy (A) loses valuable data (survivor narratives are historically underrepresented)
  2. Public access is wrong (clear privacy violation)
  3. Restricted access balances values:
  4. Preserve for future (historical value)
  5. Protect privacy (researchers must apply, IRB oversight)
  6. Allow opt-out (if users emerge, they can request removal)
  7. Seal (B) or Research Access (C) depends on:
  8. How identifiable are users? (If highly identifiable → seal)
  9. How urgent is research need? (If active crisis → research access)

If choosing C (research access):

  • Require IRB approval
  • Redact identifying info (usernames, locations, specific details)
  • Content warnings
  • Share only with trauma-informed researchers
  • Partner with domestic violence organizations (they can advise on safety)

Lesson: When in doubt, preserve but restrict. You can always open access later (with community input), but you can’t un-lose destroyed data.


Part IV: Building Institutional Ethics Frameworks §

Creating a Preservation Ethics Policy §

If you’re building an archive or preservation organization, codify your ethical approach:

1. Mission and Values Statement

Example: “We preserve LGBTQ+ digital culture to ensure queer histories are not erased. We prioritize:

  • Community consent and self-determination
  • Marginalized voices over mainstream narratives
  • Harm reduction and privacy protection
  • Transparency in our preservation decisions”

2. Preservation Criteria

What will you preserve?

  • Cultural significance (how defined?)
  • Community connection (must be created by/for LGBTQ+ people?)
  • Time period (all eras, or specific focus?)
  • Content types (text, images, video, all of the above?)

How do you handle consent?

  • Ideal: Explicit consent from creators
  • Pragmatic: Implied consent for public content, with opt-out
  • Restricted: Preserve sensitive material with limited access
  • Never: No consent-violating leaks, private messages, stolen data

4. Access Tiers

Who can see what?

Tier 1: Public Access

  • Fully public material, no privacy concerns
  • Searchable, downloadable

Tier 2: Researcher Access

  • Apply for access, state research purpose
  • IRB approval if studying human subjects

Tier 3: Community Access Only

  • Only LGBTQ+ researchers/community members
  • Protects community autonomy

Tier 4: Sealed

  • Preserved but not accessible (yet)
  • Time embargo (open in X years)

5. Takedown and Appeal Process

How do people request removal?

  • Submit request via form
  • Review by ethics committee
  • Decision within 30 days
  • Appeal process if denied

Automatic takedown for:

  • Non-consensual intimate images
  • Doxxing (personal addresses, phone numbers)
  • Minors depicted (if requestor is the minor, now adult)

6. Ethical Review Board

Who makes hard decisions?

  • Staff members + community advisors + ethicists
  • Meets monthly to review contested cases
  • Decisions documented and published (redacted for privacy)

Example: Internet Archive’s Approach (Simplified)

Mission: Universal access to knowledge

Consent: Respect robots.txt (if site owner says “don’t crawl,” they don’t)

Access: Public by default

Takedown: Submit DMCA or personal information removal request, reviewed within days

Strengths:

  • Simple, clear
  • Respects technical consent signals (robots.txt)
  • Fast takedown process

Limitations:

  • Doesn’t proactively consider harm
  • Public default may violate privacy norms
  • Legal framework (DMCA) ≠ ethical framework

Better model would add:

  • Ethics board for ambiguous cases
  • Proactive harm assessment (don’t wait for takedown requests)
  • Tiered access for sensitive material

Part V: Personal Ethics for Individual Archivists §

Your Own Custodial Ethics §

If you’re preserving as an individual (not an institution), you still need ethical frameworks:

1. Know Your Biases

Reflect:

  • What do I think is important? (And what am I overlooking?)
  • Whose voices am I centered? (Am I reproducing mainstream biases?)
  • What communities do I have connections to? (And which am I an outsider to?)

Action:

  • Actively seek out marginalized perspectives
  • Defer to community members on what matters
  • Recognize your limitations

Steps:

  • If preserving someone’s work, try to contact them
  • If contact is impossible, assume consent for public content but allow opt-out
  • If content is borderline (public but privacy-sensitive), err on side of restriction

3. Document Your Decisions

Keep a log:

  • What did you preserve and why?
  • What did you skip and why?
  • What ethical dilemmas arose?
  • How did you resolve them?

Reasons:

  • Transparency (others can critique your choices)
  • Learning (you’ll improve over time by reviewing past decisions)
  • Accountability (if someone challenges you, you have reasoning)

4. Seek Input

Don’t decide alone:

  • Join communities of practice (Archive Team, preservation forums)
  • Ask others for advice on hard cases
  • Accept that you’ll make mistakes, be open to correction

5. Provide Escape Hatches

Build in ways for people to undo your preservation:

  • Public contact info (email, form) for takedown requests
  • Honor requests promptly and without judgment
  • Apologize when you get it wrong

Part VI: When Preservation Is Wrong §

Artifacts That Should Not Be Preserved §

Some content should be actively destroyed, not preserved:

1. Child Sexual Abuse Material (CSAM)

Never preserve. Full stop.

  • Illegal
  • Victimizes children
  • No historical or research value that justifies harm

If you encounter: Report to NCMEC (National Center for Missing & Exploited Children), delete immediately.

2. Non-Consensual Intimate Images (NCII / “Revenge Porn”)

Never preserve.

  • Severe privacy violation
  • Ongoing harm to victims
  • Criminal in many jurisdictions

If you encounter: Delete, report to platform or law enforcement if active

3. Doxxing That Endangers Lives

Do not preserve if:

  • Content includes personal addresses, phone numbers
  • Clear intent to enable harassment or violence
  • Victim is at active risk

Exception: If part of larger newsworthy event (Jan 6 insurrection), redact personal info but preserve rest

4. Terrorist Manifestos with Actionable Plans

Do not preserve if:

  • Detailed instructions for violence
  • Clear intent to inspire copycat attacks
  • Ongoing threat

Exception: Preserve metadata (that it existed, summary) without full content. Work with law enforcement if actively dangerous.

5. Content Explicitly Illegal in Your Jurisdiction

Know the law:

  • Some countries ban Holocaust denial, hate speech
  • US has broader speech protections (but CSAM, true threats are still illegal)

If in doubt: Consult lawyer before preserving


Conclusion: The Weight of the Gavel §

When you preserve, you wield power. You decide what future generations remember. You shape who gets voice and who gets silence. You determine what harms are perpetuated and what accountability is enabled.

This is not a burden to take lightly.

The Custodial Filter gives you structure for these decisions—five questions that force ethical deliberation:

  1. Does this matter? (Cultural significance)
  2. Do people consent? (Consent)
  3. Does preserving cause harm? (Harm assessment)
  4. Is someone else doing this? (Redundancy)
  5. Can I realistically do this? (Feasibility)

But structure isn’t certainty. You will face dilemmas where all options feel wrong. You will make mistakes. You will carry the weight of artifacts you couldn’t save and artifacts you shouldn’t have saved.

This is the custodial burden.

Accept it. Feel it. Let it make you careful. But don’t let it paralyze you.

Because the alternative—doing nothing—is also an ethical choice. And when platforms murder culture, doing nothing is complicity.

So preserve. But preserve ethically. Save what matters, respect those who don’t want to be saved, restrict access when harm is possible, document your reasoning, and build escape hatches.

Be a custodian who wields power with humility, who makes hard choices transparently, and who accepts accountability for the consequences.

The gavel is heavy. Carry it anyway.

In the next chapter, we explore the Triage Workflow—the eight-step process for moving from discovery to preservation to access. Now that we understand the ethics, we’ll learn the systematic methodology.


Discussion Questions §

  1. The Survivor Forum: What would you have done with the domestic abuse survivor forum hard drive? Justify your choice using the Custodial Filter.

  2. Consent Boundaries: Where do you draw the line? At what point does historical significance override individual desire to be forgotten?

  3. Personal Bias: What are your own biases about what’s “important” to preserve? How do those biases shape what gets remembered?

  4. Harm Weighing: How do you weigh potential future research value against present harm? Can you quantify that trade-off?

  5. Institutional vs. Individual: Should ethics be different for large institutions (Internet Archive) vs. individual archivists? Why or why not?

  6. The Custodial Burden: Have you ever made a preservation decision? How did it feel? What did you learn?


Exercise: Ethical Triage Simulation §

Scenario: You’re part of an archival collective. A platform announces shutdown in 72 hours. Your group has capacity to preserve 20% of the platform’s content. Here are six types of content—rank them 1-6 for preservation priority, then justify your ranking:

Content Types:

A. Celebrity Accounts (1 million posts)

  • High engagement, widely discussed
  • Already screenshotted/quoted in media
  • Creators are wealthy, can hire archivists if they want

B. LGBTQ+ Youth Support Forum (50,000 posts)

  • Private/semi-private discussions
  • Coming out stories, mental health support
  • Many posters were minors at time of posting
  • No other known archive

C. Political Misinformation Archive (200,000 posts)

  • Conspiracy theories, false health info
  • Harmful but historically significant
  • Researchers studying radicalization want access

D. Fan Fiction Community (500,000 stories)

  • Transformative works, LGBTQ+ representation
  • Some authors deleted stories (wanted them gone)
  • Copyright gray area (derivative of copyrighted works)

E. Indie Artist Portfolios (100,000 works)

  • Original art, music, poetry
  • Many artists no longer active online (can’t contact)
  • Some nsfw content (but artistic, not pornographic)

F. Corporate Brand Accounts (500,000 posts)

  • Marketing, customer service interactions
  • Low cultural value but documents commercial strategies
  • Easy to scrape (public, well-structured)

Part 1: Ranking (300 words)

  • Rank 1-6 (1 = highest priority)
  • Justify each ranking using the five Custodial Filter questions

Part 2: Ethical Dilemmas (500 words)

For your top 2 choices, identify:

  • What ethical dilemmas arise?
  • How would you handle consent?
  • What access restrictions (if any)?
  • How would you handle takedown requests?

Part 3: Reflection (300 words)

  • What was hardest about ranking?
  • Did you prioritize harm reduction, cultural significance, or feasibility?
  • How did your personal values shape your choices?
  • Would you make different choices if you had more time/resources?

Further Reading §

On Archival Ethics §

  • Caswell, Michelle. Urgent Archives: Enacting Liberatory Memory Work. Routledge, 2021.
  • Jimerson, Randall. Archives Power: Memory, Accountability, and Social Justice. Society of American Archivists, 2009.
  • Flinn, Andrew. “Community Histories, Community Archives: Some Opportunities and Challenges.” Journal of the Society of Archivists 28, no. 2 (2007): 151-176.
  • Nissenbaum, Helen. Privacy in Context: Technology, Policy, and the Integrity of Social Life. Stanford University Press, 2009.
  • Boyd, Danah. “Privacy and Publicity in the Context of Big Data.” WWW ‘14 Keynote, 2014.
  • Marwick, Alice, and danah boyd. “Networked privacy: How teenagers negotiate context in social media.” New Media & Society 16, no. 7 (2014): 1051-1067.

On Harm and Trauma-Informed Practice §

  • Caswell, Michelle, et al. “‘To Be Able to Imagine Otherwise’: Community Archives and the Importance of Representation.” Archives and Records 38, no. 1 (2017): 5-26.
  • Herman, Judith. Trauma and Recovery. Basic Books, 1992.
  • Substance Abuse and Mental Health Services Administration. SAMHSA’s Concept of Trauma and Guidance for a Trauma-Informed Approach. 2014.

On Digital Ethics §

  • Markham, Annette, and Elizabeth Buchanan. Ethical Decision-Making and Internet Research. Association of Internet Researchers, 2012.
  • Zimmer, Michael. “But the data is already public: on the ethics of research in Facebook.” Ethics and Information Technology 12, no. 4 (2010): 313-325.

End of Chapter 9

Next: Chapter 10 — Triage Workflow: From Discovery to Preservation

Part II • Excavation Methods & Digital Forensics

Chapter 10: Triage Workflow

From Discovery to Preservation

24 min read 5,104 words

Opening: The Clock Is Always Ticking §

March 17, 2023, 9:47 AM: A Discord message in the Archive Team channel: “Credit Karma is shutting down their forums on April 15th. 28 days. Thousands of posts about personal finance from 2007-2023. Anyone on this?”

9:52 AM: Three people respond. They’ve never worked together before. One is a college student in California. One is a librarian in Germany. One is a retired programmer in Ohio.

10:15 AM: They’ve created a shared spreadsheet, assigned tasks, and started reconnaissance.

April 14th, 11:58 PM: The scraping is complete. 47,000 posts, 8,200 users, 16 years of financial advice—all captured. Total time: 27 days, 14 hours. They did it.

April 15th, 12:01 AM: Credit Karma’s forums go offline. The original URLs return 404 errors. But the archive exists—backed up to Internet Archive, stored on three personal servers, uploaded as a torrent.

This is triage workflow in action: from discovery to preservation in less than a month. Every step matters. Every hour counts. One mistake, one delay, and the content is lost forever.

This chapter teaches you the complete triage workflow—an 8-phase process tested across hundreds of platform deaths. Whether you have 48 hours or 6 months, this framework will guide you from panic to preservation.


The 8-Phase Triage Workflow §

Overview §

Phase 1: Discovery — Detecting that content is endangered
Phase 2: Assessment — Understanding scope, urgency, and feasibility
Phase 3: Mobilization — Assembling team and resources
Phase 4: Capture — Executing the scrape/download/preservation
Phase 5: Validation — Verifying data integrity
Phase 6: Storage — Securing long-term preservation
Phase 7: Access — Making content discoverable and usable
Phase 8: Documentation — Recording what you did and why

Each phase has specific goals, tools, and decision points. Let’s explore them in detail.


Phase 1: Discovery — Detecting Endangerment §

Goal §

Identify that content is at risk of disappearing before it’s too late.

Common Discovery Channels §

1. Official Announcements

  • Platform posts shutdown notice (GeoCities, Vine, Google+)
  • Company blog, email to users, platform notification
  • Timeline: Usually 30-90 days warning (sometimes less)

2. Financial/Business Signals

  • Company files bankruptcy
  • Acquisition by competitor (often precedes shutdown)
  • Mass layoffs, especially engineering
  • Stopping development (no updates in 12+ months)
  • Timeline: Months to years before actual shutdown

3. User Exodus

  • Mass migration to alternatives
  • “I’m leaving [platform], find me at…” posts
  • Decline in active users
  • Timeline: Can indicate slow death (years) or precede rapid collapse

4. Technical Degradation

  • Frequent outages
  • Bugs not being fixed
  • Security vulnerabilities left unpatched
  • Timeline: Months before shutdown (or years of zombie state)

5. Community Monitoring

  • Archive Team’s Deathwatch (tracks endangered sites)
  • Social media warnings (Twitter/Reddit threads)
  • Journalism (tech news covering potential shutdowns)
  • Timeline: Varies (can be early warning or last-minute)

6. Policy Changes

  • Terms of Service updates that hostile to users
  • Monetization changes (introducing paywalls, removing features)
  • Content purges (Tumblr NSFW ban)
  • Timeline: Immediate (policy goes into effect) or weeks

Discovery Tools and Practices §

Proactive Monitoring:

  • Subscribe to platform announcements (email, RSS, social media)
  • Use website monitoring tools (detect when site goes down)
  • Follow tech journalism (The Verge, TechCrunch, Ars Technica)
  • Join Archive Team Discord/IRC (community shares warnings)

Reactive Response:

  • When you hear rumor, investigate immediately
  • Don’t wait for official confirmation (sometimes never comes)
  • Err on side of caution (preserve early rather than late)

Decision Point: Is This Worth Investigating? §

Rapid Assessment (5 minutes):

  • Scale: How much content exists?
  • Cultural value: Does anyone care if this disappears?
  • Urgency: How imminent is the threat?
  • Existing preservation: Is someone else already handling this?

If answers suggest “yes, endangered and valuable,” proceed to Phase 2.


Phase 2: Assessment — Understanding the Challenge §

Goal §

Determine scope, technical requirements, ethical concerns, and resource needs before committing to preservation.

Assessment Checklist §

A. Scope Assessment

Content Inventory:

  • How many pages/posts/users/files?
  • What media types? (text, images, video, audio, documents)
  • What time span? (1 year? 20 years?)
  • What languages/communities represented?

Example: Credit Karma Forums

  • 47,000 posts across 12 subforums
  • Text + occasional images (hosted externally)
  • 2007-2023 (16 years)
  • Primarily English, US-focused

Storage Estimate:

  • Text: ~500 words/post × 47k posts = 23.5M words ≈ 150MB text
  • Images: ~500 external links, assume 20% capturable = 100 images × 500KB = 50MB
  • Total estimated: 200MB (small! very doable)

B. Technical Assessment

Platform Architecture:

  • Static HTML or dynamic JavaScript?
  • Public access or login-required?
  • API available? Rate limits?
  • Search/browse mechanisms?

Preservation Difficulty:

  • Easy (wget-able): ★☆☆☆☆
  • Medium (requires browser automation): ★★★☆☆
  • Hard (heavily DRM’d, real-time only): ★★★★★

Tools Needed:

  • Basic scraping: wget, HTTrack, ArchiveBox
  • Dynamic sites: Selenium, Playwright, browser automation
  • API harvesting: Python scripts, API clients
  • Forensic recovery: specialized tools (if site partially dead)

Example: Credit Karma Forums Technical Profile

  • Dynamic site (JavaScript-rendered pagination)
  • Login required (but free account creation)
  • No public API
  • Difficulty: ★★★☆☆ (need browser automation + account)

C. Urgency Assessment

Time Until Loss:

  • Shutdown announced: Count down from announcement date
  • No announcement but signs of death: Estimate (weeks? months?)
  • Already partially dead: URGENT (capture what remains)

Timeline Categories:

  • Critical (< 1 week): Drop everything, act now
  • Urgent (1-4 weeks): High priority, mobilize quickly
  • High (1-3 months): Important, plan thoroughly
  • Medium (3-6 months): Time for systematic approach
  • Low (6+ months): Monitor, begin planning

Example: Credit Karma

  • 28 days from discovery to shutdown = Urgent
  • Can’t be leisurely, but can plan

D. Resource Assessment

Labor:

  • Can one person do this? Or need team?
  • How many hours estimated?
  • What skills needed? (coding, systems admin, metadata, etc.)

Infrastructure:

  • How much storage? (Do you have it?)
  • Bandwidth? (Will download take days?)
  • Computing power? (Scraping 10M pages needs beefy machine)

Budget:

  • Free (volunteer labor + personal resources)?
  • Small budget ($100-1000 for servers/storage)?
  • Grant-funded ($10k+ for major project)?

Example: Credit Karma

  • Labor: 2-3 people, ~40 hours each (part-time over 4 weeks)
  • Storage: 200MB (trivial—USB drive sufficient)
  • Bandwidth: Minimal (small text files)
  • Budget: $0 (volunteer effort)

E. Ethical Assessment (Custodial Filter)

Cultural Significance: Medium-high (personal finance advice, especially recession-era)

Technical Fragility: High (28 days to shutdown)

Rescue Feasibility: Medium (doable with browser automation)

Redundancy: None (no other known preservation effort)

Ethical Concerns:

  • Privacy: Posts may contain personal financial details
  • Consent: Users didn’t expect permanent archiving
  • Harm potential: Low (financial advice, not doxxing or harassment)

Decision: Preserve with restricted access

  • Capture everything
  • Researcher-only access (not public searchable web)
  • Allow user-requested takedowns

Copyright:

  • Who owns the content? (Platform TOS usually claims license, but users retain copyright)
  • Fair use argument? (Archiving for research/scholarship)
  • DMCA risk? (Platform could issue takedown if they notice)

Terms of Service:

  • Does TOS forbid scraping? (Usually yes, but rarely enforced for preservation)
  • Are you violating contract by scraping? (Technically yes, but ethical override)

Privacy Laws:

  • GDPR (if EU users)? Right to be forgotten vs. archival interest
  • CCPA (California)? Data export vs. data retention

Risk Assessment:

  • Low risk: Defunct platform unlikely to sue preservationists
  • Medium risk: Active platform might send cease-and-desist
  • High risk: Legally protected content (DRM, government secrets)

Example: Credit Karma

  • Fair use: Strong argument (educational/research archiving)
  • TOS violation: Yes, but platform dying (unlikely to enforce)
  • Privacy: Medium concern (financial discussions)
  • Risk: Low overall (proceed but don’t publicize widely)

Output of Phase 2: Go/No-Go Decision §

After assessment, decide:

GO: Proceed with preservation

  • Timeline: [realistic schedule]
  • Team needed: [number of people, skills]
  • Tools required: [specific software/hardware]
  • Budget: [if any]
  • Ethical framework: [access restrictions, takedown policy]

NO-GO: Don’t preserve (because…)

  • Too large (beyond capacity)
  • Too technically difficult (lack skills/tools)
  • Ethically problematic (more harm than good)
  • Redundant (someone else already doing it better)
  • Not urgent (can wait, revisit later)

DEFER: Monitor but don’t act yet

  • Not urgent enough
  • Waiting for more information
  • Hoping platform survives

Phase 3: Mobilization — Assembling Resources §

Goal §

Get team, tools, and infrastructure ready before capture begins.

3A: Team Formation §

Solo vs. Collaborative:

When to work solo:

  • Small project (< 100 hours work)
  • Simple tools (basic scraping)
  • No deadline pressure (can take months)

When to recruit team:

  • Large project (> 100 hours)
  • Tight deadline (need parallel effort)
  • Specialized skills needed (you can’t do everything)

Recruiting:

  • Archive Team Discord/IRC: Post call for volunteers
  • Social media: Twitter, Reddit (r/DataHoarder)
  • Academic networks: Colleagues, students
  • Local communities: Library listservs, tech meetups

Team Roles:

  • Coordinator: Manages overall effort, tracks progress
  • Technical lead: Designs scraping strategy, writes code
  • Scrapers: Run tools, troubleshoot issues
  • Validators: Check data integrity, spot gaps
  • Metadata curator: Organizes captured content
  • Legal/ethical advisor: Navigates consent/privacy issues

Example: Credit Karma Team

  • 3 volunteers (found via Archive Team)
  • Coordinator = college student (had free time, organized)
  • Technical lead = retired programmer (wrote scraping scripts)
  • Scraper = librarian (ran tools, captured pages)

3B: Tool Selection and Setup §

Scraping Tools:

Static Sites:

  • wget (command-line, recursive downloading)
  • HTTrack (GUI, mirrors entire websites)
  • ArchiveBox (modern, all-in-one archiving)

Dynamic Sites (JavaScript-heavy):

  • Selenium (browser automation, Python/Java)
  • Playwright (modern alternative to Selenium)
  • Puppeteer (Node.js browser control)

API Harvesting:

  • PRAW (Reddit API, Python)
  • Tweepy (Twitter API, Python)
  • Custom scripts (platform-specific APIs)

Forensic Recovery:

  • Webrecorder (captures dynamic content, WARC format)
  • Heritrix (Internet Archive’s crawler)
  • Browsertrix Crawler (cloud-based crawling)

Example: Credit Karma Stack

  • Selenium (browser automation for dynamic pagination)
  • Python (scripting)
  • SQLite (local database to track progress)
  • rsync (backup to multiple locations)

3C: Infrastructure Setup §

Storage:

  • Local: External hard drives (cheap, reliable for small projects)
  • Cloud: AWS S3, Google Cloud Storage (for large projects)
  • Distributed: IPFS, BitTorrent (censorship-resistant)
  • Institutional: University servers, Internet Archive

Compute:

  • Personal laptop (small projects)
  • VPS (DigitalOcean, Linode) (medium projects, avoid IP bans)
  • Cloud compute (AWS EC2) (large-scale scraping)

Bandwidth:

  • Residential internet: Usually sufficient (but may hit caps)
  • VPS/cloud: Unmetered bandwidth (expensive but fast)

Backup Strategy:

  • 3-2-1 rule: 3 copies, 2 different media types, 1 offsite
  • Real-time sync (rsync while scraping, don’t wait until end)
  • Checksums (verify data integrity)

Example: Credit Karma Infrastructure

  • Storage: 3 USB drives (1 per team member) + Internet Archive upload
  • Compute: Personal laptops (no VPS needed, small scale)
  • Bandwidth: Residential (200MB download didn’t strain anything)

3D: Coordination Tools §

Communication:

  • Discord/Slack (real-time chat)
  • GitHub Issues (track tasks, bugs)
  • Shared spreadsheet (who’s doing what, progress tracking)

Documentation:

  • Wiki or shared doc (technical notes, scraping strategies)
  • Git repository (for code)
  • Progress log (daily updates)

Example: Credit Karma Coordination

  • Discord private channel (3-person team)
  • Google Spreadsheet (tracking forum sections, who scraped what)
  • GitHub repo (Python scripts + documentation)

Phase 4: Capture — Executing the Preservation §

Goal §

Download/scrape/capture the endangered content before it disappears.

4A: Capture Strategy §

Breadth vs. Depth:

  • Breadth-first: Capture as many items as possible (may be shallow)
  • Depth-first: Capture complete item details (may miss some items)

Example:

  • Breadth: Scrape all post titles and links (fast, ensures you have IDs)
  • Depth: Download full post content, comments, attachments (slower, but complete)

Best practice: Breadth first (get IDs of everything), then depth (fill in details). If time runs out, you at least have a list of what existed.

Parallelization:

  • Multiple scrapers running simultaneously (different sections/users)
  • Be careful: Too aggressive = IP ban
  • Use delays, rotate IPs/user agents

4B: Capture Execution §

Step 1: Initial Crawl (Breadth)

  • Map the site structure (what sections exist?)
  • Enumerate all items (posts, users, pages)
  • Store URLs/IDs in database

Step 2: Content Download (Depth)

  • For each item, fetch full content
  • Save text, images, videos, metadata
  • Record relationships (replies, quotes, etc.)

Step 3: Iterative Refinement

  • Identify gaps (missing content, broken links)
  • Re-scrape failed items
  • Validate as you go (don’t wait until end)

Example: Credit Karma Capture Process

Day 1-3: Reconnaissance

  • Manually browse forums to understand structure
  • Identify 12 subforums, ~4,000 threads per forum
  • Estimate: 47,000 posts total

Day 4-7: Initial Crawl

  • Selenium script navigates forum pages
  • Extracts thread IDs and post IDs
  • Stores in SQLite database (47,211 post IDs captured)

Day 8-20: Content Download

  • For each post ID, fetch:
  • Post text (HTML + plaintext)
  • Author username and join date
  • Timestamp (posted date)
  • Like/reply counts
  • Quoted text (if reply)
  • Save as JSON files (one per post)
  • Progress: ~2,500 posts/day (3 people × ~800 posts each)

Day 21-26: Gap Filling

  • Identified 342 posts that failed to download (timeouts, errors)
  • Re-scraped with slower rate
  • Success: 47,155 / 47,211 (99.88% capture rate)

Day 27: Final Validation

  • Spot-checked 100 random posts (all looked good)
  • Calculated checksums for all files
  • Created manifest (list of all files + checksums)

4C: Dealing with Technical Challenges §

Challenge 1: Rate Limiting

  • Platform blocks you after X requests per minute
  • Solution: Add delays (time.sleep() between requests), rotate IPs (VPN/proxies), use multiple accounts

Challenge 2: JavaScript Rendering

  • Content doesn’t appear in raw HTML (loaded by JS)
  • Solution: Use browser automation (Selenium, Playwright), not simple wget

Challenge 3: Login Walls

  • Need account to access content
  • Solution: Create throwaway account (use privacy-respecting email), automate login in scraper

Challenge 4: CAPTCHAs

  • Bot detection blocks automated access
  • Solution: Manual CAPTCHA solving, CAPTCHA-solving services (2captcha, anti-captcha), reduce scraping speed (look more human)

Challenge 5: Dynamic URLs

  • URLs change on each visit (session IDs, tokens)
  • Solution: Extract content IDs, construct stable URLs, use API if available

Challenge 6: Server Instability

  • Dying platform’s servers are flaky (timeouts, errors)
  • Solution: Retry failed requests, save progress frequently, accept imperfect capture

4D: Ethical Boundaries During Capture §

Don’t:

  • Overload dying servers (cause outage for remaining users)
  • Scrape private content without consent
  • Violate clear legal restrictions (DRM-protected content)
  • Ignore takedown requests (if someone asks you to stop, consider it)

Do:

  • Be respectful (slow scraping, don’t hammer servers)
  • Document decisions (why you preserved X but not Y)
  • Provide opt-out (let people request removal later)

Phase 5: Validation — Verifying Data Integrity §

Goal §

Ensure captured data is complete, accurate, and uncorrupted.

5A: Completeness Checks §

Quantitative:

  • Did you get everything? (Compare scraped count to expected count)
  • How many items failed? (Acceptable loss rate: < 1%)

Example: Credit Karma

  • Expected: 47,211 posts
  • Captured: 47,155 posts
  • Missing: 56 posts (0.12% loss—acceptable)

Qualitative:

  • Random sampling: Spot-check 100 items (do they look correct?)
  • Edge cases: Check first post, last post, longest post, shortest post
  • Relationships: Do replies correctly link to parent posts?

5B: Integrity Checks §

File Corruption:

  • Generate checksums (MD5, SHA-256) for every file
  • Verify checksums after transfer (detect corruption during copy)

Format Validation:

  • Are files readable? (Open JSON, parse HTML, view images)
  • Are encodings correct? (UTF-8 for text, not garbled)

Metadata Accuracy:

  • Timestamps make sense? (No posts “from the future”)
  • Usernames consistent? (No missing or duplicated users)

5C: Documentation of Gaps §

What’s Missing:

  • List items you couldn’t capture (with reasons)
  • Example: “Posts 234, 457, 891 returned 404 (already deleted before scrape)”

Known Issues:

  • Broken images (external links dead)
  • Incomplete threads (some replies missing)
  • Corrupted formatting (HTML parser issues)

Why Documentation Matters:

  • Future researchers need to know limits of collection
  • Transparency about what’s incomplete
  • Legal protection (“we preserved what we could access”)

Phase 6: Storage — Long-Term Preservation §

Goal §

Store captured data securely with redundancy for decades-long access.

6A: Storage Formats §

Raw Captures:

  • WARC files (Web ARChive format—standard for web preservation)
  • JSON (structured data, easy to parse)
  • Database dumps (SQL exports)

Derived Formats:

  • Static HTML (browseable offline)
  • PDFs (human-readable, archival quality)
  • CSV (for datasets, spreadsheet-compatible)

Media:

  • Images: PNG/JPEG (lossless or high-quality)
  • Video: MP4/WebM (widely supported codecs)
  • Audio: FLAC/MP3 (lossless or high-bitrate)

Example: Credit Karma Storage

  • Primary: JSON files (one per post) + SQLite database
  • Derived: Static HTML site (browseable offline)
  • Uploaded: WARC files to Internet Archive

6B: Redundancy Strategy §

Local Redundancy:

  • Multiple hard drives (3+ copies)
  • Different physical locations (not all in one apartment)

Cloud Redundancy:

  • Upload to Internet Archive (free, stable institution)
  • AWS S3 Glacier (cheap long-term storage, but not free)
  • IPFS (distributed, censorship-resistant)

Community Redundancy:

  • Torrent (BitTorrent allows distributed hosting)
  • Share with other archivists (trust network)

Example: Credit Karma Redundancy

  1. USB drive #1 (coordinator’s backup)
  2. USB drive #2 (technical lead’s backup)
  3. USB drive #3 (scraper’s backup)
  4. Internet Archive upload (public institution)
  5. BitTorrent (uploaded to Archive Team tracker)

Result: 5 copies, multiple custodians, extremely unlikely to be fully lost

6C: Metadata Preservation §

Collection-level metadata:

  • What is this? (Credit Karma Forums archive)
  • When was it captured? (March-April 2023)
  • Who captured it? (Archive Team volunteers)
  • How much? (47,155 posts, 8,200 users)
  • What’s missing? (56 posts unavailable)

Item-level metadata:

  • Post ID, author, timestamp, content, replies, likes

Technical metadata:

  • Scraping tools used
  • Checksums for files
  • File formats and encodings

Store metadata in:

  • README.md (human-readable)
  • JSON manifest (machine-readable)
  • Dublin Core XML (library standard)

Phase 7: Access — Making Content Discoverable §

Goal §

Ensure captured content is usable by researchers, communities, and the public.

7A: Access Levels §

Public Access:

  • Full content searchable on open web
  • Appropriate for: Public posts, low privacy concerns, culturally significant
  • Example: Internet Archive’s Wayback Machine

Researcher Access:

  • Require application, academic credentials, or agreement
  • Appropriate for: Sensitive content, privacy concerns, contested material
  • Example: University archives with restricted collections

Community Access:

  • Available only to original community members
  • Appropriate for: Private forums, cultural protocols (Indigenous archives)
  • Example: Invite-only Discord with archived content

Dark Archive:

  • Preserved but not accessible (yet)
  • Appropriate for: Ethically fraught content, legal uncertainty, time embargo
  • Example: Preserve now, review access in 50 years

Example: Credit Karma Access Decision

  • Public metadata (list of posts, authors, dates)—searchable
  • Full content researcher-only (financial details sensitive)
  • User-requested takedowns honored

7B: Access Infrastructure §

Static HTML Site:

  • Browse/search offline or on local server
  • Tools: Custom scripts, static site generators (Hugo, Jekyll)

Database + Web Interface:

  • MySQL/PostgreSQL + web app (Flask, Django)
  • Allows advanced search, filtering

Upload to Platforms:

  • Internet Archive (free hosting, discoverable)
  • GitHub (for code + datasets < 100GB)
  • Zenodo (academic datasets, DOI assignment)

Example: Credit Karma Access

  • Static site generated from JSON (browseable offline)
  • Uploaded to Internet Archive (public metadata + researcher-access content)
  • Torrent (full download for other archivists)

7C: Discovery Mechanisms §

How do people find this archive?

Documentation:

  • Blog post announcing completion
  • Social media (Twitter, Reddit)
  • Archive Team wiki page

Indexing:

  • Internet Archive indexed by Google
  • Scholarly databases (if deposited to university)

Community Outreach:

  • Contact original users (if possible)
  • Notify journalists/researchers who might care

Phase 8: Documentation — Recording the Process §

Goal §

Document what you did, why, and what happened for future archivists and researchers.

8A: Technical Documentation §

Scraping Process:

  • What tools? What settings?
  • How long did it take?
  • What worked? What failed?

Challenges Encountered:

  • Technical problems and solutions
  • Rate limiting workarounds
  • Data quality issues

Final Statistics:

  • Items captured
  • Storage size
  • Time spent
  • Team size

Example: Credit Karma Documentation (excerpt)

bash
# Credit Karma Forums Archive - Technical Documentation

## Timeline
- Discovery: March 17, 2023
- Capture: March 19 - April 14, 2023 (27 days)
- Shutdown: April 15, 2023

## Team
- 3 volunteers (Archive Team)

## Tools
- Selenium (Python) for browser automation
- SQLite for progress tracking
- rsync for backups

## Statistics
- 47,155 posts captured (99.88% of estimated total)
- 8,200 unique users
- 2007-2023 (16 years of content)
- Total size: 187 MB (compressed)

## Challenges
- Dynamic pagination required browser automation
- Server timeouts during peak hours (scraped during US nighttime)
- 56 posts returned 404 (likely deleted by users before scrape)

## Storage
- 5 redundant copies (3 USB drives, Internet Archive, BitTorrent)

8B: Ethical Documentation §

Decisions Made:

  • Why this content?
  • Why this access level?
  • How did you handle privacy/consent?

Takedown Policy:

  • How can people request removal?
  • What’s your response process?

Future Considerations:

  • Should access restrictions be lifted eventually?
  • Who decides?

8C: Historical Documentation §

Why This Mattered:

  • What was this platform’s cultural significance?
  • Who used it? For what?
  • Why did it die?

Contextual Essay:

  • Write 500-1000 words explaining the platform and its role
  • Include for future researchers who won’t remember it

Example: Credit Karma Context (excerpt)

Credit Karma was a free credit-monitoring service that launched forums in 2007. During the Great Recession (2008-2009), these forums became a vital resource for people navigating financial hardship—debt, bankruptcy, foreclosure, unemployment. Users shared advice, support, and strategies for rebuilding credit. The forums remained active through 2023, documenting 16 years of American financial struggles and recovery. Credit Karma shut down the forums as part of a platform redesign focused on mobile apps. The decision prioritized sleek user experience over community memory, erasing nearly two decades of peer support and financial education.

8D: Lessons Learned §

What Would You Do Differently?

  • Start earlier?
  • Use different tools?
  • Recruit more help?

Advice for Future Archivists:

  • What worked well?
  • What to avoid?

Meta-Reflection:

  • How did this project change your thinking about preservation?

Case Study: The Complete Triage Workflow in Action §

The GeoCities Rescue (2009) §

Let’s trace the entire workflow through Archive Team’s legendary GeoCities rescue:

Phase 1: Discovery

  • October 26, 2009: Yahoo announces GeoCities will shut down November 26
  • Archive Team hears via tech news
  • Timeline: 31 days

Phase 2: Assessment

  • Scope: ~30 million sites (estimated 10+ TB of data)
  • Technical: Simple HTML (mostly), but massive scale
  • Urgency: Critical (1 month)
  • Resources: Volunteer network (Archive Team has ~50 active members)
  • Ethics: Public content, cultural treasure, no privacy concerns
  • Decision: GO (this is huge, we must try)

Phase 3: Mobilization

  • Team: ~100 volunteers recruited (IRC, Twitter)
  • Roles: Coordinators (tracked progress), scrapers (ran wget), validators (checked captures)
  • Tools: wget (simple but effective for static HTML)
  • Infrastructure: Personal computers + Internet Archive upload
  • Coordination: Archive Team IRC channel (24/7 chat)

Phase 4: Capture

  • Strategy: Divide by “neighborhoods” (GeoCities organized sites into themed sections: /SiliconValley/, /Tokyo/, etc.)
  • Each volunteer took a neighborhood
  • Breadth-first: Enumerate all sites, then download content
  • Challenges: Yahoo rate-limited aggressive scrapers (volunteers rotated IPs, added delays)
  • Result: 650 GB captured (fraction of total, but significant)

Phase 5: Validation

  • Checked file counts, looked for corruption
  • Known gaps: Many sites missed (not enough time/bandwidth)
  • Completeness: ~10-15% (still, 650GB is massive)

Phase 6: Storage

  • Uploaded to Internet Archive
  • Created BitTorrent (distributed to hundreds of archivists)
  • Multiple volunteers kept local copies

Phase 7: Access

  • Internet Archive hosts browseable interface
  • BitTorrent available for full download
  • OoCities.org (community project to host GeoCities mirrors)

Phase 8: Documentation

  • Archive Team wiki documents entire process
  • Media coverage (Wired, Ars Technica, NPR)
  • Academic papers cite GeoCities archive

Legacy:

  • GeoCities rescue established Archive Team’s reputation
  • Proved that volunteer networks can save platforms
  • Inspired future rescues (Vine, Google+, Yahoo Groups, etc.)

Workflow Variations for Different Scenarios §

Scenario 1: Emergency Triage (< 48 hours) §

Compress the workflow:

  • Phase 1 (Discovery): Immediate (someone alerts you)
  • Phase 2 (Assessment): 30 minutes (quick check)
  • Phase 3 (Mobilization): 1 hour (grab tools, alert others)
  • Phase 4 (Capture): 24-36 hours (aggressive scraping, accept gaps)
  • Phase 5 (Validation): Minimal (spot-check, full validation later)
  • Phase 6 (Storage): Quick upload to Internet Archive
  • Phase 7 (Access): Defer (just preserve first, organize later)
  • Phase 8 (Documentation): Brief notes (expand later)

Priority: Speed over perfection. Save something rather than nothing.

Scenario 2: Systematic Preservation (6+ months) §

Expand the workflow:

  • Phase 2: Thorough assessment (weeks of analysis)
  • Phase 3: Professional team (hire researchers, not just volunteers)
  • Phase 4: Careful capture (high fidelity, complete metadata)
  • Phase 5: Extensive validation (manual review of samples)
  • Phase 6: Archival-quality storage (institutional partnerships)
  • Phase 7: Polished access (custom web interface, finding aids)
  • Phase 8: Scholarly publication (write paper on process + findings)

Priority: Quality and comprehensiveness. Create gold-standard archive.

Stealth considerations:

  • Phase 3: Work solo or small trusted team (no public announcements)
  • Phase 4: Use VPN/Tor (anonymize scraping)
  • Phase 6: Encrypted storage (protect yourself if content controversial)
  • Phase 7: Dark archive or anonymous torrent (not publicly credited)
  • Phase 8: Minimal documentation (protect participants)

Priority: Survival (yours and the archive’s). Preserve ethically but carefully.


Conclusion: The Workflow Is Your Map §

Digital preservation under deadline is chaos. Platforms die with little warning. Servers vanish. URLs break. The clock ticks down.

The workflow is your map through chaos. It won’t make preservation easy, but it will make it systematic. When you panic (and you will), return to the workflow:

  1. What phase am I in?
  2. What’s the goal of this phase?
  3. What’s the next action?

The workflow has been tested across hundreds of platform deaths. It works. It’s saved millions of digital artifacts. It will guide you through your first rescue—and your hundredth.

Next chapter: Part III: Institution Building begins. We’ve learned to excavate, analyze, triage, and preserve. Now we must build institutions that can sustain this work for decades—organizations that outlive founders, survive funding crises, and resist corporate capture.

The rescue is only the beginning. The real work is building systems that prevent future murders.

But first: practice the workflow. Find an endangered platform. Walk through the phases. Preserve something.

The clock is always ticking. Start now.


Discussion Questions §

  1. Personal Experience: Have you ever tried to preserve digital content before a deadline (even personal—like backing up your own social media)? What went well? What did you wish you’d known?

  2. Workflow Adaptation: Which scenario (emergency, systematic, guerrilla) would be hardest for you? Why? What skills would you need to develop?

  3. Team Dynamics: The Credit Karma example had 3 strangers collaborate effectively. What made that work? What could go wrong?

  4. Validation Trade-offs: In emergency triage, validation is minimal. How do you decide “good enough” when perfection isn’t possible?

  5. Access Decisions: The Credit Karma archive restricted full content to researchers. Agree or disagree? Where would you draw the line?

  6. Future Scenarios: Imagine a platform shutdown in 2030. What might be different (technology, laws, culture)? How would the workflow need to adapt?


Exercise: Conduct a Practice Triage §

Scenario: It’s November 2025. A small platform called “BookTalk” (fictional) announces it will shut down in 60 days. It’s a reading discussion forum with:

  • 5,000 users
  • 250,000 posts (book reviews, discussion threads)
  • 15 years of history (2010-2025)
  • Dynamic JavaScript site (requires login)

Your Task: Walk through the 8-phase workflow.

Phase 1: Discovery (Already done—you just heard the news)

Phase 2: Assessment (500 words)

  • Complete the assessment checklist
  • Scope, technical, urgency, resources, ethics, legal
  • Make a GO/NO-GO decision

Phase 3: Mobilization (300 words)

  • Would you work solo or recruit a team?
  • What tools would you use?
  • What infrastructure do you need?

Phase 4: Capture (500 words)

  • Design your scraping strategy
  • Breadth vs. depth approach
  • Timeline (how many days for each step?)
  • What could go wrong?

Phase 5: Validation (200 words)

  • How would you verify completeness?
  • What checks would you run?

Phase 6: Storage (300 words)

  • What formats?
  • How many redundant copies?
  • Where would you store them?

Phase 7: Access (300 words)

  • What access level? (public, researcher, community, dark)
  • Why?
  • How would people discover this archive?

Phase 8: Documentation (200 words)

  • What would you document?
  • For whom?

Reflection (300 words)

  • What was hardest to decide?
  • What would you need to learn to actually do this?
  • Would you commit to this project? Why or why not?

Further Reading §

On Preservation Workflows §

  • Bailey, Jefferson. “Disrespect des Fonds: Rethinking Arrangement and Description in Born-Digital Archives.” Archive Journal 3 (2013).
  • Lee, Christopher. “A Framework for Contextual Information in Digital Collections.” Journal of Documentation 67, no. 1 (2011): 95-143.
  • Prom, Christopher. “Managing Risks in Web Archiving: Best Practices and Guidelines.” Digital Preservation Coalition, 2018.

On Rapid Response Archiving §

  • Archive Team. “Warrior Documentation.” https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior
  • Brügger, Niels. “Web Archiving: The Urgent Need for Preservation.” In Web Archiving, edited by Niels Brügger and Ralph Schroeder, 1-20. MIT Press, 2017.
  • Summers, Ed, et al. “Learning to Crawl: Towards a Framework for the History of Web Archiving.” International Journal of Digital Humanities 1 (2019): 105-124.

On Data Integrity and Validation §

  • Duranti, Luciana, and Corinne Rogers. “Trust in Digital Records: An Increasingly Cloudy Legal Area.” Computer Law & Security Review 28, no. 5 (2012): 522-531.
  • Rosenthal, David S.H. “Formats Considered Harmful.” iPRES 2014 Conference (2014).

On Access and Ethics §

  • Caswell, Michelle. “Seeing Yourself in History: Community Archives and the Fight Against Symbolic Annihilation.” The Public Historian 36, no. 4 (2014): 26-37.
  • Punzalan, Ricardo, and Michelle Caswell. “Critical Directions for Archival Approaches to Social Justice.” Library Quarterly 86, no. 1 (2016): 25-42.

Primary Sources §

  • Archive Team Wiki. https://wiki.archiveteam.org/
  • Internet Archive. “Wayback Machine Documentation.” https://archive.org/web/
  • Digital Preservation Coalition. “Rapid Assessment Model.” https://www.dpconline.org/

End of Chapter 10

Next: Part III — Institution Building Chapter 11 — Sustainable Preservation Organizations: Building the Archive

Part III • Institutional Design, Commons & Political Economy

Chapter 11: Sustainable Preservation Organizations

Building the Archive

21 min read 4,618 words

Opening: The 50-Year Question §

In 1996, Brewster Kahle founded the Internet Archive with a simple but audacious mission: “Universal access to all knowledge.” Nearly 30 years later, it’s still here—800+ billion web pages archived, 40+ million books scanned, millions of videos and audio recordings preserved.

But walk through the graveyard of digital preservation projects that didn’t make it:

  • Google’s various archiving experiments (Google+, Google Reader, countless other shutdowns)
  • Yahoo’s acquisitions (GeoCities, Delicious—bought then killed)
  • University projects that folded when the PhD student graduated or the grant ended
  • Volunteer projects that died when the founder burned out

The brutal truth: Most preservation organizations fail. They launch with enthusiasm, run for a few years, then collapse when funding dries up, founders leave, or technology becomes obsolete.

This chapter asks: How do you build a preservation organization that survives 50 years?

Not 5 years (easy with grant funding). Not 10 years (possible with dedicated volunteers). But 50+ years—long enough to outlive founders, survive technological shifts, weather economic recessions, and fulfill the actual promise of “long-term” preservation.

We’ll examine the Internet Archive as a successful model, analyze why most preservation organizations fail, and provide a framework for designing institutions that can endure.


Part I: Why Most Preservation Organizations Fail §

Failure Mode 1: The Heroic Founder Problem §

The Pattern:

  • Charismatic founder with vision and energy launches preservation project
  • Founder works 80-hour weeks, sustaining the project through sheer will
  • Organization is shaped around founder’s skills, networks, and personality
  • Founder eventually leaves (burnout, new job, death)
  • Organization collapses because it can’t function without founder

Examples:

Aaron Swartz and RECAP (Successful Succession)

  • Aaron created RECAP to scrape PACER (US court documents) and make them freely available
  • When Aaron died (2013), the project could have died with him
  • But: He’d built it with institutional partners (Princeton’s CITP)
  • Survived transition because it wasn’t dependent on one person

Countless volunteer archives (Failed Succession)

  • Fan archives, forum backups, small community preservation projects
  • Typically run by one dedicated volunteer
  • When that person stops (school, job, burnout), archive goes offline
  • No succession plan, no institutional memory

Why It Happens:

  • Founders are often “lone wolves”—prefer doing to delegating
  • Building governance structures feels like bureaucracy (slows you down)
  • Ego: “I know how to do this best”
  • Practical: Hard to find others with same skills/commitment

How to Prevent:

  • Distribute leadership early (co-founders, boards, committees)
  • Document everything (operations manuals, decision logs, technical docs)
  • Build governance structures before crisis (succession plan, transfer protocols)
  • Train deputies (people who can step in if founder leaves)

Failure Mode 2: Financial Fragility §

The Pattern:

  • Organization launches with initial funding (grant, donation, venture capital)
  • Operates well during funded period
  • Funding runs out (grant expires, donors lose interest, VCs demand profit)
  • Organization shuts down or becomes extractive (paywalls, surveillance, ads)

Examples:

Delicious (VC Capture → Death)

  • Popular bookmarking service (founded 2003)
  • Acquired by Yahoo (2005) then sold to others multiple times
  • Each new owner tried to monetize, failed
  • Slowly decayed, users abandoned it
  • Now zombie platform (technically alive but culturally dead)

Pinboard (Sustainable Model)

  • Bookmarking service launched 2009 as Delicious alternative
  • Paid model: One-time fee ($11) then yearly archiving fee ($25/year)
  • No ads, no surveillance, no investor pressure
  • Profitable with one developer (Maciej Cegłowski)
  • Still running 15+ years later

Why It Happens:

  • Grants are temporary (2-5 years typically)
  • Donations are volatile (depend on donor whims, economic conditions)
  • VC funding demands exits (IPO or acquisition), corrupts mission
  • Free services must monetize eventually (ads/surveillance)

How to Prevent:

  • Diversify funding (multiple sources, not dependent on one grant/donor)
  • Earned revenue (services, subscriptions, bulk sales to institutions)
  • Endowments (build capital that generates interest)
  • Non-profit structure (legally protected from acquisition/profit pressure)
  • Frugality (keep costs low, avoid growth addiction)

Failure Mode 3: Technical Obsolescence §

The Pattern:

  • Organization builds preservation infrastructure on current technology
  • Technology evolves (storage formats change, APIs deprecate, protocols shift)
  • Organization can’t keep up with technical maintenance
  • Preserved content becomes inaccessible (bit rot, format obsolescence, broken systems)

Examples:

Early CD-ROM Archives (Format Death)

  • 1990s: Libraries burned archives to CD-ROMs (cutting-edge preservation)
  • 2020s: CD drives rare, many CDs degraded, formats obsolete
  • Content preserved but inaccessible

Floppy Disk Archives (Media Decay)

  • Museums have floppy disks with important software/data
  • Disks demagnetized over time
  • Drives increasingly rare
  • Many archives lost because couldn’t be read in time

Why It Happens:

  • Preservation = long-term, but technology = short-term (5-10 year cycles)
  • Format migration is expensive (labor-intensive)
  • “Set it and forget it” doesn’t work (requires continuous maintenance)

How to Prevent:

  • Open formats (avoid proprietary; use standards like WARC, XML, JSON)
  • Format migration budget (allocate money/time to periodic updates)
  • Redundancy (multiple copies in multiple formats/locations)
  • Emulation (preserve both data AND tools to read it)
  • Active monitoring (check integrity regularly, don’t assume it’s fine)

Failure Mode 4: Scope Creep and Mission Drift §

The Pattern:

  • Organization starts with clear, focused mission
  • Success attracts new opportunities (grants for related work, partnerships, expansions)
  • Organization takes on too many projects
  • Resources spread thin, core mission neglected, eventually collapses

Examples:

Internet Archive (Avoided This)

  • Could have expanded into dozens of areas (cloud storage, social media platform, publishing, etc.)
  • Instead: Stayed focused on core mission (web archiving, book scanning, media preservation)
  • Does new projects (e.g., lending library) but only if aligned with mission

Many University DH Centers (Fell Into This)

  • Start as DH research centers
  • Get asked to support all tech needs (“Can you help with faculty websites?” “Can you run our server?”)
  • Become IT support, lose research focus
  • Funding cut because they’re not producing research

Why It Happens:

  • Hard to say no (money, opportunities, helping people)
  • Mission creep gradual (each small expansion seems reasonable)
  • No clear boundaries (if you preserve web, why not apps? If apps, why not games? If games, why not…”)

How to Prevent:

  • Written mission statement (refer to it when deciding new projects)
  • Clear boundaries (what you do, what you don’t do)
  • Sunset projects (end things that don’t serve mission, even if popular)
  • Focus metrics (measure success by depth, not breadth)

Failure Mode 5: Community Disconnection §

The Pattern:

  • Organization preserves content but doesn’t engage community it serves
  • Community doesn’t feel ownership or investment
  • When crisis comes (funding cut, founder leaves), no one fights to save it

Examples:

Academic Archives (Often This Problem)

  • Universities build archives of local history
  • Store them in basements, minimal access
  • Community doesn’t know they exist
  • When budget cuts come, archives closed (no public outcry)

Wikipedia (Avoided This)

  • Millions of contributors feel ownership
  • When funding threatened, massive support campaigns
  • Community would fight to keep it alive

Why It Happens:

  • Preservation seen as expert/institutional work (not community involvement)
  • Access restricted (researchers only, not public)
  • No communication (organization doesn’t share what it’s doing)

How to Prevent:

  • Public engagement (regular updates, blog posts, social media)
  • Open access (make archives usable, not locked in vaults)
  • Community involvement (volunteers, advisory boards, user-generated metadata)
  • Visible impact (show how archives are used, who benefits)

Part II: The Internet Archive Model — What Works §

Why the Internet Archive Has Survived 30 Years §

Let’s analyze what makes Internet Archive successful as a blueprint for other organizations:

1. Clear, Compelling Mission

“Universal access to all knowledge.”

  • Simple enough for anyone to understand
  • Ambitious enough to inspire
  • Specific enough to guide decisions (does this project advance universal access?)
  • Defensible in court (when sued, mission provides moral/legal argument)

Lesson: Your mission should be:

  • Memorable (one sentence)
  • Motivating (people want to support it)
  • Actionable (guides what you do/don’t do)

2. Non-Profit Structure with Diverse Funding

Legal Structure:

  • 501(c)(3) non-profit (US tax-exempt status)
  • Can’t be acquired by for-profit companies
  • Donations tax-deductible (incentivizes giving)

Revenue Sources (diversified):

  • Individual donations ($20-$100 from millions of users)
  • Foundation grants (Mellon, Sloan, Knight, etc.)
  • Earned revenue (digitization services for libraries)
  • Government contracts (scanning books for Library of Congress)
  • Partnerships (library consortia paying for services)

Why Diversification Matters:

  • If one funding source disappears, others sustain operations
  • Not beholden to any single funder’s whims
  • Can maintain independence (no investor pressure)

Lesson: Aim for at least 3 major revenue streams, none exceeding 50% of budget.

3. Technical Excellence with Open Standards

Technology Choices:

  • WARC format (web archives, open standard)
  • Open-source software (Heritrix crawler, Wayback Machine code)
  • Commodity hardware (not expensive proprietary systems)
  • Distributed architecture (not single point of failure)

Why This Matters:

  • Open standards mean others can replicate/extend their work
  • Open source builds community (others contribute code)
  • Commodity hardware keeps costs down (can scale cheaply)
  • Distributed systems survive disasters (data centers can fail without losing everything)

Lesson: Build on open standards, avoid lock-in, design for redundancy.

4. Institutional Partnerships

Key Partnerships:

  • Library of Alexandria (mirror site in Egypt)
  • 1,000+ libraries worldwide (distributed preservation network)
  • Universities (research collaborations, storage nodes)
  • Foundations (funders who understand long-term mission)

Why Partnerships Matter:

  • Redundancy (if Internet Archive burns down, partners have copies)
  • Legitimacy (respected institutions endorse the work)
  • Resources (partners contribute storage, bandwidth, expertise)
  • Resilience (distributed network harder to destroy)

Lesson: Don’t go alone. Build alliances with institutions that share your values.

5. Public Engagement and Transparency

Engagement Strategies:

  • Public-facing website (anyone can browse archives)
  • Blog (regular updates on projects, challenges, victories)
  • Media appearances (Brewster Kahle gives talks, interviews)
  • Open financials (publishes annual reports, 990 tax forms)
  • Appeals for support (fundraising drives with clear goals)

Why This Matters:

  • Users feel invested (they know what’s happening)
  • Transparency builds trust (donors see where money goes)
  • Public support protects from political threats (when attacked, people defend you)

Lesson: Communicate constantly. Make your work visible. Build constituency.

Legal Work:

  • Fights copyright battles (defended right to lend digital books)
  • Joined Brewster Kahle’s personal activism (right to repair, DRM opposition)
  • Files amicus briefs in tech cases
  • Part of broader movement (EFF, Public Knowledge, etc.)

Why This Matters:

  • Preservation exists in legal gray areas (fair use, copyright exceptions)
  • Without legal advocacy, vulnerable to shutdown
  • Fighting back sets precedents (protects other preservation organizations)

Lesson: Budget for legal work. Join coalitions. Fight bad laws.

7. Founder Who Built Succession

Brewster Kahle’s Approach:

  • Hired strong leadership team (not one-person show)
  • Created board of directors (governance beyond himself)
  • Diversified expertise (librarians, technologists, lawyers, fundraisers)
  • Built institution, not personal fiefdom

Result: If Brewster left tomorrow, Internet Archive would survive (though it’d be hard). It’s not dependent on him alone.

Lesson: Build succession from day one, even when it feels premature.


Part III: The Archive Business Canvas §

To design a sustainable preservation organization, use this framework:

Section 1: Mission and Values §

Core Mission (one sentence):

  • What are you preserving? Why?
  • Example: “Preserve murdered social media platforms to document 21st-century culture”

Core Values (3-5 principles):

  • What guides your decisions?
  • Example: “User sovereignty, comprehensive preservation, open access, community accountability, ethical triage”

Success Metrics (how you know you’re succeeding):

  • Quantitative: TB preserved, items cataloged, users served
  • Qualitative: Community trust, cultural impact, policy influence

Section 2: What You Preserve §

Scope (be specific):

  • What artifacts? (websites, videos, games, software, etc.)
  • What time period? (everything since 1990? last 10 years?)
  • What geographic/cultural focus? (global? specific communities?)

Triage Criteria:

  • Cultural significance thresholds
  • What you explicitly DON’T preserve (ethical boundaries)

Curation Philosophy:

  • Comprehensive (save everything you can) vs. selective (curate carefully)
  • How much metadata/context do you add?

Section 3: Technical Infrastructure §

Storage:

  • How much capacity needed? (current + growth projections)
  • Where stored? (cloud, local servers, distributed network)
  • Redundancy? (how many copies, where?)

Access Systems:

  • How do people find/use preserved content?
  • Web interface, APIs, physical access?
  • Public vs. restricted access?

Preservation Methods:

  • Formats used (WARC, JPEG, MP4, etc.)
  • Emulation needs (for complex artifacts)
  • Format migration schedule (how often you update)

Technical Team:

  • Who builds/maintains systems?
  • In-house developers? Contractors? Volunteers?

Section 4: Governance §

Legal Structure:

  • Non-profit (501(c)(3) in US, charity elsewhere)? For-profit social enterprise? Cooperative?
  • Why this structure? What does it protect against?

Leadership:

  • Who makes decisions? (board, director, community vote?)
  • How are leaders chosen? (election, appointment, rotation?)
  • Term limits? (prevent capture by entrenched leadership)

Succession Plan:

  • What happens if founder leaves?
  • Who takes over? How is transition managed?

Accountability:

  • How do you ensure mission fidelity?
  • Community oversight? Board review? External audits?

Section 5: Funding Model §

Revenue Streams (diversify):

Donations:

  • Individual (small donors via website)
  • Major donors (wealthy individuals/foundations)
  • Membership (recurring monthly/yearly)

Earned Revenue:

  • Services (digitization for institutions, consulting)
  • Bulk sales (datasets to researchers, corporations)
  • Licensing (allowing commercial use for fee)

Grants:

  • Government (NEH, IMLS, NSF)
  • Foundations (Mellon, Sloan, Knight, Mozilla)
  • Universities (research collaborations)

Partnerships:

  • Institutions pay for shared infrastructure
  • Collaborative grants (multiple organizations)

10-Year Budget Projection:

  • Years 1-2: Grant-dependent (foundation seed funding)
  • Years 3-5: Diversify (add earned revenue, donations)
  • Years 6-10: Sustainable mix (no single source >40%)
  • Years 10+: Endowment building (create permanent capital)

Section 6: Staffing §

Core Team (full-time):

  • Executive Director (leadership, fundraising, strategy)
  • Technical Director (systems architecture, engineering)
  • Curatorial Lead (metadata, collection development, access)
  • Development Director (fundraising, donor relations)

Extended Team (part-time or contractors):

  • Legal counsel
  • Grant writer
  • Developers (for specific projects)
  • Metadata specialists
  • Community manager

Volunteers:

  • What roles can volunteers fill? (metadata tagging, quality assurance, community outreach)
  • How do you recruit, train, retain?

Growth Plan:

  • Start small (2-3 people)
  • Add roles as funding grows
  • Don’t hire until revenue sustains salary long-term

Section 7: Community and Partnerships §

Primary Community:

  • Who creates the content you preserve?
  • Who uses your archives?
  • How do you engage them?

Advisory Board:

  • Representatives from community
  • Guide priorities, review decisions
  • Provide legitimacy and accountability

Institutional Partners:

  • Libraries, universities, museums
  • What do they contribute? What do they get?
  • Formal agreements or informal collaborations?

Coalitions:

  • Join existing networks (NDSA, DPC, etc.)
  • Advocacy coalitions (fight bad laws together)
  • Technical collaborations (share tools, standards)

Section 8: Risk Management §

Threats:

Financial:

  • Major donor withdraws → Mitigation: Diversified funding
  • Grant rejected → Mitigation: Multiple applications, earned revenue buffer

Technical:

  • Data center fire → Mitigation: Geographic redundancy, backups
  • Cyber attack → Mitigation: Security audits, offline backups

Legal:

  • Copyright lawsuit → Mitigation: Legal fund, insurance, coalition support
  • Government shutdown order → Mitigation: International mirrors, legal advocacy

Organizational:

  • Founder burnout → Mitigation: Succession plan, distributed leadership
  • Staff turnover → Mitigation: Documentation, redundant expertise

Reputational:

  • Scandal (preserved harmful content) → Mitigation: Ethics policy, transparency, community accountability

Existential:

  • Mission becomes irrelevant (e.g., problem solved) → Mitigation: Broad mission, adaptable strategy

Section 9: Three Pillars Integrity Check §

Declaration:

  • Do you own your infrastructure? (Domain, servers, code)
  • Can you operate independently of platforms?

Connection:

  • Do you engage community directly?
  • Can you communicate without corporate intermediaries?

Ground:

  • Do you own your data and tools?
  • Can you migrate if hosting providers fail?

If any Pillar is weak, strengthen it before launch.


Part IV: Three Preservation Organization Models §

Model 1: The Institutional Archive (Internet Archive Model) §

Characteristics:

  • Large-scale, comprehensive preservation
  • Professional staff (10-100+ employees)
  • $5M-$50M annual budget
  • Non-profit, foundation-funded + earned revenue
  • Decades-long time horizon

Best For:

  • National/international scope
  • Multiple content types
  • Serving researchers and public
  • Long-term stability

Example Adaptations:

  • Regional Internet Archive (state or country-specific)
  • Subject-specific (gaming archive, music archive, etc.)

Model 2: The Community Archive (Fan Archive Model) §

Characteristics:

  • Small-scale, focused collection
  • Volunteer-run (0-3 paid staff)
  • $10K-$100K annual budget
  • Donations + community funding
  • Community ownership/governance

Best For:

  • Specific communities (fandoms, subcultures, local history)
  • Grassroots preservation
  • High community trust needed

Examples:

  • Archive of Our Own (fan fiction)
  • Small museum archives (local historical societies)
  • Discord/forum archives (community-run)

Challenges:

  • Sustainability (volunteers burn out)
  • Succession (what happens when founder leaves?)
  • Funding (hard to get grants without institutional backing)

Mitigations:

  • Rotate leadership (share burden)
  • Partner with institution (university hosts/backs you)
  • Federate (join network of similar archives)

Model 3: The Cooperative Archive (Distributed Ownership Model) §

Characteristics:

  • Member-owned (users/creators collectively own it)
  • Democratic governance (one member, one vote)
  • Revenue from members ($10-$100/year per person)
  • Federated infrastructure (multiple nodes)

Best For:

  • Communities that want sovereignty
  • Platforms with existing user base
  • Avoiding corporate capture

Examples:

  • Stocksy (photographer cooperative)
  • Resonate (musician cooperative)
  • Mastodon instances (community-run servers)

Challenges:

  • Coordination (democratic = slower decisions)
  • Technical (members must understand governance)
  • Scale (hard to grow quickly)

Strengths:

  • Resilient (distributed ownership)
  • Aligned incentives (users are also owners)
  • Mission-protected (can’t be sold out)

Part V: Launching Your Preservation Organization §

Phase 1: Planning (6-12 months before launch) §

Step 1: Clarify Mission

  • Write one-sentence mission
  • Define 3-5 core values
  • Set specific scope (what you preserve, what you don’t)

Step 2: Assess Feasibility

  • Is there demand? (do people want this?)
  • Is it sustainable? (can you fund it for 10+ years?)
  • Are you the right team? (do you have necessary skills?)

Step 3: Build Core Team

  • Find co-founders (don’t go alone)
  • Recruit advisors (people with relevant expertise)
  • Form initial board (if non-profit)

Step 4: Prototype

  • Preserve a small sample (prove you can do it)
  • Build minimal access system (test usability)
  • Get feedback (from potential users)

Step 5: Secure Seed Funding

  • Apply for initial grants (Mellon, NEH, Knight)
  • Recruit anchor donors (wealthy individuals who believe in mission)
  • Estimate 2-year runway (enough to launch and prove viability)

Phase 2: Launch (Months 1-12) §

Step 1: Legal Formation

  • Incorporate (as non-profit, cooperative, or other structure)
  • Apply for tax-exempt status (if non-profit)
  • Set up banking, accounting

Step 2: Build Infrastructure

  • Secure storage (servers, cloud, or hybrid)
  • Build access systems (website, search, API)
  • Set up preservation workflows

Step 3: Preserve Initial Collections

  • Start with highest-priority content
  • Document methods (create operations manual)
  • Add metadata and context

Step 4: Engage Community

  • Launch public website
  • Announce on social media, press
  • Recruit volunteers, advisors, users

Step 5: Prove Sustainability

  • Diversify funding (don’t rely on one grant)
  • Track metrics (preservation volume, user engagement, financial health)
  • Iterate based on feedback

Phase 3: Growth (Years 2-5) §

Step 1: Scale Operations

  • Hire staff (as funding allows)
  • Expand storage and preservation capacity
  • Improve access systems

Step 2: Build Partnerships

  • Join preservation networks (NDSA, DPC)
  • Partner with institutions (libraries, museums)
  • Collaborate with similar organizations

Step 3: Develop Earned Revenue

  • Offer services (digitization, consulting, bulk data)
  • Launch membership program
  • Reduce grant dependency

Step 4: Invest in Governance

  • Formalize decision-making processes
  • Create succession plans
  • Build board capacity

Step 5: Communicate Impact

  • Publish annual reports
  • Share success stories
  • Demonstrate value to funders and community

Phase 4: Maturity (Years 5+) §

Step 1: Build Endowment

  • Capital campaign (raise money for permanent fund)
  • Interest sustains base operations
  • Reduces financial volatility

Step 2: Ensure Succession

  • Executive transitions (leadership changes smoothly)
  • Institutional memory preserved
  • Independence from any one person

Step 3: Expand Impact

  • Influence policy (testify, draft legislation)
  • Train next generation (offer internships, workshops)
  • Export model (help others start similar organizations)

Step 4: Maintain Mission

  • Regular reviews (are you still serving original purpose?)
  • Resist scope creep
  • Sunset projects that don’t fit

Step 5: Plan for Century

  • What happens in 50 years?
  • How will technology change?
  • How do you ensure continuity?

Part VI: Case Studies of Successful Archives §

Case Study 1: Archive of Our Own (AO3) §

Context:

  • Fan fiction archive (user-written stories based on existing media)
  • Founded 2008 by Organization for Transformative Works (OTW)
  • Response to FanFiction.Net’s increasing restrictions and commercial archive shutdowns

Model:

  • 501(c)(3) non-profit
  • Volunteer-run (100+ volunteers, no paid staff for archive itself)
  • Donation-funded ($200K-$300K annual budget from small donors)
  • No ads, no data mining, no corporate ownership

Key Success Factors:

  • Community ownership: Users feel it’s “theirs”
  • Clear values: Pro-fan, anti-censorship, anti-commercial
  • Distributed volunteer team: No single point of failure
  • Legal backing: OTW provides legal support (copyright defense)

Challenges:

  • Volunteer burnout (high turnover)
  • Scaling (millions of works, storage costs growing)
  • Governance (democratic but can be slow)

Lessons:

  • Community archives can work at scale if values align
  • Volunteer models require strong succession/rotation
  • Clear mission attracts passionate contributors

Case Study 2: Software Heritage §

Context:

  • Archive of all public source code (GitHub, GitLab, etc.)
  • Founded 2016 by INRIA (French research institute)
  • Mission: Preserve “software commons” for future

Model:

  • Non-profit foundation
  • Government/foundation grants + institutional partnerships
  • €2M-€5M annual budget
  • Professional staff + academic researchers

Key Success Factors:

  • Institutional backing: INRIA provides stability
  • UNESCO partnership: Cultural heritage recognition
  • Technical excellence: Advanced deduplication, graph structures
  • Open data: Anyone can access archives

Challenges:

  • Scope (40+ million projects and growing)
  • Legal ambiguity (copyright on archived code unclear)
  • Sustainability (still grant-dependent after 8 years)

Lessons:

  • Institutional backing accelerates legitimacy
  • Technical innovation can be competitive advantage
  • Even well-funded projects face sustainability challenges

Case Study 3: Perma.cc §

Context:

  • Preserves links cited in legal and scholarly documents
  • Founded 2013 by Harvard Library Innovation Lab
  • Prevents “link rot” in citations

Model:

  • University-backed (hosted by Harvard)
  • Hybrid funding (free for individuals, subscriptions for institutions)
  • Small team (2-3 people)
  • Narrow, well-defined scope

Key Success Factors:

  • Focused mission: One problem, solved deeply
  • Institutional integration: Adopted by law reviews, journals
  • Sustainability: Small budget, university-backed, subscription revenue

Challenges:

  • Limited scope (only preserves cited links, not comprehensive)
  • Dependent on Harvard (if university cuts funding, vulnerable)

Lessons:

  • Narrow scope can be strength (easier to sustain, clear value proposition)
  • University backing provides stability but creates dependency
  • Hybrid funding (free + paid) balances access and sustainability

Conclusion: Building to Last §

The graveyard of digital preservation projects is vast. Most fail within 5 years. But it doesn’t have to be this way.

The keys to building archives that last:

  1. Mission clarity (know what you’re preserving and why)
  2. Financial diversity (never depend on one funding source)
  3. Technical openness (use standards, avoid lock-in, design for migration)
  4. Community engagement (build constituency that will fight for you)
  5. Governance beyond founders (succession plans, distributed leadership)
  6. Legal resilience (budget for fights, join coalitions, advocate for policy)
  7. Realistic scope (focus deeply, resist mission creep)

The Internet Archive is 30 years old. It could live another 50—or 100. Not because Brewster Kahle is superhuman, but because he built an institution, not a project.

You can do the same. Whether you’re building a massive institutional archive, a small community collection, or a cooperative platform, the principles are the same:

Design for decades. Plan for crisis. Build to outlast yourself.

In the next chapter, we’ll turn to the Anvil—the practice of building platforms and tools that embody sovereignty. Archives preserve the past; Anvils forge the future. But both require the same institutional discipline: building things that last.


Discussion Questions §

  1. Failure Analysis: Which failure mode (heroic founder, financial fragility, technical obsolescence, scope creep, community disconnection) seems most dangerous? Why?

  2. Internet Archive Sustainability: Could Internet Archive survive without Brewster Kahle? What would change? What vulnerabilities remain?

  3. Funding Ethics: Is it ethical for preservation organizations to charge for access (paywalls)? Where’s the line between sustainability and extracting value from public goods?

  4. Community vs. Institution: Should archives be community-run (like AO3) or institutionally-backed (like Software Heritage)? Trade-offs?

  5. Your Own Organization: If you were launching a preservation organization, which model (institutional/community/cooperative) would you choose? Why?

  6. 50-Year Horizon: What threats to long-term sustainability do you think we’re not considering? What will matter in 2075 that we can’t predict now?


Exercise: Design Your Preservation Organization §

Task: Design a preservation organization using the Archive Business Canvas.

Scenario: Choose one of:

  • (A) Archive for a specific murdered platform (e.g., Vine, Tumblr NSFW, Google+)
  • (B) Archive for a living but endangered community (e.g., small forums, indie web, etc.)
  • (C) Federated preservation network (distributed across institutions)

Part 1: Mission and Scope (500 words)

  • One-sentence mission
  • 3-5 core values
  • What you preserve (specific artifacts, time period, cultural focus)
  • What you explicitly DON’T preserve

Part 2: Technical Infrastructure (500 words)

  • Storage strategy (how much, where, redundancy)
  • Access systems (how users find/use content)
  • Preservation methods (formats, emulation, migration schedule)

Part 3: Governance and Staffing (500 words)

  • Legal structure (non-profit, cooperative, etc.)
  • Leadership (who decides, how chosen)
  • Succession plan
  • Team composition (who you hire, when)

Part 4: Funding Model (500 words)

  • 3+ revenue streams
  • 10-year budget projection
  • Path to sustainability
  • Risk mitigation (what if major funder withdraws?)

Part 5: Community Engagement (500 words)

  • Who are your primary users/community?
  • How do you engage them?
  • Advisory structures
  • Partnerships

Part 6: Risk Assessment (500 words)

  • Top 5 threats to your organization
  • Mitigation strategies for each
  • “What would kill us?” analysis

Part 7: Three Pillars Check (300 words)

  • Does your organization embody Declaration, Connection, Ground?
  • Any weaknesses? How do you address them?

Further Reading §

On Organizational Sustainability §

  • Ostrom, Elinor. Governing the Commons. Cambridge University Press, 1990.
  • Principles for sustaining shared resources long-term

  • Benkler, Yochai. The Wealth of Networks. Yale University Press, 2006.

  • Economics of peer production and commons-based organizations

  • Bollier, David. Think Like a Commoner. New Society Publishers, 2014.

  • Accessible introduction to commons governance

On Non-Profit Management §

  • Worth, Michael J. Nonprofit Management: Principles and Practice. 5th ed. SAGE, 2019.
  • Practical guide to running non-profits

  • Kim, Sojung, and Mason, David. “Governance and Accountability in Nonprofit Organizations.” In The Jossey-Bass Handbook of Nonprofit Leadership and Management, edited by David Renz, 317-345. Wiley, 2016.

On Digital Preservation Institutions §

  • Lavoie, Brian, and Lorcan Dempsey. “Thirteen Ways of Looking at…Digital Preservation.” D-Lib Magazine 10, no. 7/8 (2004).
  • Strategic perspectives on preservation organizations

  • Blue Ribbon Task Force on Sustainable Digital Preservation and Access. Sustainable Economics for a Digital Planet. 2010.

  • Economics of long-term preservation

Primary Sources §

  • Internet Archive. “About the Internet Archive.” https://archive.org/about/
  • Archive of Our Own. “About the OTW.” https://www.transformativeworks.org/about/
  • Software Heritage. “About Software Heritage.” https://www.softwareheritage.org/mission/
  • Perma.cc. “About Perma.cc.” https://perma.cc/about

End of Chapter 11

Next: Chapter 12 — The Economics of Sovereignty: Building the Anvil

Part III • Institutional Design, Commons & Political Economy

Chapter 12: The Economics of Sovereignty

Building the Anvil

20 min read 4,215 words

Opening: The Business Model Trap §

In 2004, a blogger named Ev Williams sold his company Blogger to Google for an undisclosed sum (reportedly millions). Blogger had pioneered user-friendly blogging, giving millions of people their own publishing platforms. But it never found a sustainable business model. Google acquired it, integrated it with their ad network, and kept it alive—but blogger.com users became Google users, subject to Google’s terms, surveillance, and whims.

In 2013, Williams tried again with Medium. This time, he’d learned: start with a business model. Medium launched with subscriptions, then experimented with advertising, then memberships, then partnerships. Each pivot changed what Medium was—from open platform to paywalled publication network to algorithmic recommendation engine. Writers never knew if their URLs would persist, if their audiences belonged to them, or if Medium would exist in five years.

In 2023, Medium still exists—but so many writers have left (fed up with pivots and VC pressure) that it’s become a shadow of its early promise. The writers who stayed are tenants, not sovereigns.

The Anvil faces a brutal question: How do you build a business that embodies digital sovereignty—one that gives users Declaration, Connection, and Ground—while also generating enough revenue to survive?

This isn’t just a technical problem. It’s an economic design problem. And it’s perhaps the hardest challenge Archaeobytology faces: forging sustainable alternatives to platform capitalism.

This chapter explores:

  • Why traditional business models fail sovereignty (VC funding, advertising, data extraction)
  • Alternative revenue models (subscriptions, cooperatives, open core, public funding)
  • Case studies of businesses that got it right (and wrong)
  • How to design a Foundry that can survive 50 years without betraying its users

By the end, you’ll understand the economics of the Anvil—and be able to design business models that align profit with sovereignty.


Part I: Why Traditional Models Kill Sovereignty §

The Venture Capital Death Spiral §

How VC Funding Works:

  1. Startup raises money (seed round: $500k-$2M)
  2. Investors buy equity (ownership stake in company)
  3. Expectation: 10x return in 5-10 years

  4. Startup grows fast (prioritizes user growth over revenue)

  5. Burns investor money to acquire users
  6. “Growth at all costs” mentality

  7. More funding rounds (Series A, B, C: $5M, $20M, $100M+)

  8. Each round dilutes founders’ ownership
  9. Investors gain board seats, influence company direction

  10. Exit pressure (IPO or acquisition)

  11. Investors want return on investment
  12. Company must either go public (stock market) or get acquired (sold to bigger company)

  13. Monetization acceleration (squeeze users for revenue)

  14. Ads, data sales, subscription paywalls
  15. “Enshittification” (Cory Doctorow’s term): Platform degrades user experience to extract value

Why This Kills Sovereignty:

Declaration Dies:

  • To maximize ad revenue, platforms need “real names” (advertisers want demographic data)
  • Content moderation favors advertiser-friendly material (censorship of controversial but legitimate speech)
  • Platform owns user identities (can ban, suspend, or sell data without consent)

Connection Dies:

  • Algorithmic feeds prioritize “engagement” (rage-bait, controversy) over chronological connection
  • Network effects become lock-in (can’t leave because everyone’s here)
  • Platforms surveil conversations for ad targeting

Ground Dies:

  • Users don’t own their content (licensed to platform)
  • No data portability (can’t easily migrate to competitors)
  • Platform can change terms, raise prices, or shut down features without consent

Case Study: Instagram’s Enshittification

Phase 1 (2010-2012): Growth

  • Beautiful, simple photo-sharing app
  • Chronological feed
  • No ads
  • Users loved it

Phase 2 (2012): Acquisition

  • Facebook buys Instagram for $1 billion
  • Promises to keep it independent
  • Still no ads (yet)

Phase 3 (2013-2015): Monetization Begins

  • Ads introduced (2013)
  • Algorithmic feed replaces chronological (2016)
  • Users see fewer posts from friends, more from brands/influencers

Phase 4 (2016-2020): Enshittification Accelerates

  • Stories (copied from Snapchat) to keep users on platform longer
  • Reels (copied from TikTok) to compete with short video
  • Shopping features (turn platform into e-commerce)
  • Feed becomes 50% ads and recommended content, not people you follow

Phase 5 (2020-present): Users Rebel

  • Photographers and artists leave (algorithms favor video over photos)
  • “Make Instagram Instagram Again” campaign
  • But network effects trap users (can’t leave; audience is there)

Sovereignty Analysis:

  • Declaration: Users don’t own @usernames (can be suspended, handles seized)
  • Connection: Algorithm decides who sees your posts (not chronological, not transparent)
  • Ground: Content hosted on Instagram servers, no export of full quality images + metadata + social graph

Lesson: VC funding forced Facebook to extract maximum value from Instagram, degrading user experience and sovereignty.

The Advertising Model’s Incompatibility §

Why Advertising Kills Sovereignty:

1. Surveillance Becomes Necessary

  • Targeted ads require user data (behavior tracking, demographics, interests)
  • Privacy becomes impossible (every action monitored)

2. Engagement Optimization Dominates

  • Platforms optimize for “time on site” (more ads served)
  • Addictive design patterns (infinite scroll, autoplay, notifications)
  • Content moderation favors controversial material (drives engagement)

3. Algorithmic Control

  • Users can’t see chronological feeds (ads would be skipped)
  • Platforms control what you see (paid content prioritized over friends)

4. Real Name Policies

  • Advertisers want demographic certainty
  • Pseudonymity becomes impossible (violates ad targeting needs)

Case Study: Twitter/X Under Musk

Pre-Musk (2006-2022):

  • Twitter struggled with profitability (ad revenue insufficient)
  • But maintained relatively open API, chronological feed option, pseudonymity

Post-Musk (2022-present):

  • Musk buys Twitter for $44 billion (mostly debt-financed)
  • Needs to generate massive revenue to service debt
  • Results:
  • Blue checkmarks become paid ($8/month for verification)
  • API access restricted (kills third-party clients, forces users to official app with more ads)
  • Algorithmic feed becomes mandatory (can’t disable)
  • “Freedom of speech” rhetoric, but actually more censorship (of competitors, critics)

Sovereignty Impact:

  • Declaration: Verification becomes paid feature, pseudonymous users harassed
  • Connection: Third-party clients killed, API access limited, algorithmic manipulation
  • Ground: No meaningful data portability, platform instability

Lesson: Even “ideological” ownership (Musk claimed to champion free speech) gets crushed by economic pressure. Debt + ad dependence = sovereignty impossible.


Part II: Alternative Revenue Models §

If VC funding and advertising kill sovereignty, what’s left? Several models exist—none perfect, all involve trade-offs.

Model 1: Subscriptions (User-Pays) §

How It Works:

  • Users pay monthly/annual fee ($5-50/month typical)
  • Revenue funds development, infrastructure, support
  • No ads, no data sales

Sovereignty Potential: ★★★★☆

Pros:

  • No surveillance needed (users are customers, not products)
  • Incentives align (make users happy so they renew)
  • Can remain small and profitable (don’t need billions of users)

Cons:

  • Excludes people who can’t pay (equity issue)
  • Harder to grow (free platforms have network effect advantage)
  • Requires continuous value delivery (users will cancel if not worth it)

Case Study: Hey.com (Basecamp’s Email Service)

Launch (2020):

  • $99/year for email service
  • Privacy-focused (no tracking, no ads)
  • Opinionated design (built-in features for email triage)

Business Model:

  • Subscription revenue funds small team (~10 people)
  • No investors, no ads, no data sales
  • Profitable from day one (break-even at ~20,000 users)

Sovereignty Assessment:

  • Declaration: Users can use custom domains ([email protected] via Hey)
  • Connection: Email is federated (Hey users can email anyone)
  • Ground: Partial (emails stored on Hey servers, but IMAP export available)

Limitations:

  • $99/year excludes low-income users
  • Small market share (Gmail is free, Hey is not)
  • Dependent on Basecamp’s continued interest (what if they shut it down?)

Lesson: Subscriptions enable sovereignty but limit reach.

Model 2: Open Core (Free Software + Paid Hosting/Support) §

How It Works:

  • Core software is open source (free to use, modify, self-host)
  • Company offers paid hosting, support, enterprise features
  • Example: WordPress (open source) + WordPress.com (paid hosting)

Sovereignty Potential: ★★★★★

Pros:

  • Users can self-host (full sovereignty) or pay for convenience
  • Can’t be captured (if company sells out, community forks)
  • Scales (free tier grows community, paid tier funds development)

Cons:

  • Hard to compete with cloud giants (who can offer hosting cheaper)
  • Risk of “tragedy of the commons” (many users, few contributors/payers)
  • Pressure to create artificial limitations (cripple free version to upsell paid)

Case Study: Ghost (Publishing Platform)

History:

  • Founded 2013 as Kickstarter project ($300k raised)
  • Open-source blogging platform (alternative to WordPress)
  • Business model: Ghost(Pro) managed hosting + Ghost Foundation (non-profit)

Revenue Streams:

  • Ghost(Pro): $9-199/month for managed hosting
  • Self-hosting: Free (download and run yourself)
  • Foundation: Grants and donations fund open-source development

Hybrid Structure:

  • Ghost Foundation (non-profit): Owns open-source code
  • Ghost(Pro) (for-profit): Provides managed hosting, funds foundation

Sovereignty Assessment:

  • Declaration: Custom domains standard (even on free self-hosted)
  • Connection: Open API, can integrate with any service, full RSS support
  • Ground: Full ownership (self-hosted) or excellent portability (Ghost(Pro) export is complete)

Success Metrics:

  • 3,000+ paying Ghost(Pro) customers
  • Tens of thousands self-hosting
  • Profitable and sustainable (10+ years running)

Lesson: Open core + hybrid structure (non-profit + for-profit) can work.

Model 3: Cooperative Ownership (User-Owned Platform) §

How It Works:

  • Platform is owned by members (users, workers, or both)
  • Governance is democratic (one member, one vote)
  • Profits distributed to members or reinvested in platform
  • Legal structure: Co-op, worker-owned, multi-stakeholder

Sovereignty Potential: ★★★★★

Pros:

  • Structural alignment (owners are users, incentives match)
  • Can’t be sold to VCs or acquired by megacorp
  • Democratic governance (users decide platform direction)

Cons:

  • Hard to fund initial development (co-ops struggle to raise capital)
  • Governance is slow (democracy takes time)
  • Risk of capture by vocal minority (co-op politics can be messy)

Case Study: Resonate (Music Streaming Co-op)

Model:

  • Musician-owned streaming platform
  • Artists get fair pay (#stream2own: listeners “buy” songs after 9 plays)
  • Multi-stakeholder co-op (musicians, listeners, workers all have governance stake)

Funding:

  • Initial crowdfunding
  • Ongoing membership fees
  • Investment from cooperative-friendly funds

Challenges:

  • Slow growth (competing with Spotify’s billions in VC funding)
  • Technical debt (limited resources for development)
  • Governance complexity (balancing stakeholder interests is hard)

Status (2025):

  • Still operating, but small (not mainstream success)
  • Proof of concept: co-op model can work for digital platforms

Sovereignty Assessment:

  • Declaration: Artists control their profiles, own their presence
  • Connection: Direct artist-listener relationship (no algorithmic intermediation)
  • Ground: Artists own their music files, can leave platform with all data

Lesson: Co-ops align sovereignty with structure, but struggle to compete with VC-funded giants.

Model 4: Public/Non-Profit Funding (Mission-Driven) §

How It Works:

  • Platform funded by grants, donations, government funding
  • Non-profit legal structure (mission over profit)
  • Revenue sources: Philanthropic foundations (Mellon, Knight, Ford), government (NEH, NSF), individual donations

Sovereignty Potential: ★★★★☆

Pros:

  • No profit motive (mission is preservation/access, not extraction)
  • Can serve public good (not just paying customers)
  • Long time horizons (not driven by quarterly earnings)

Cons:

  • Grant dependency (what if funders change priorities?)
  • Mission drift risk (chasing grants can distort mission)
  • Slow to adapt (bureaucracy, consensus decision-making)

Case Study: Internet Archive

Funding:

  • Donations (50% of revenue): Individual donors + corporate sponsors
  • Grants (30%): Mellon Foundation, Knight Foundation, NEH
  • Services (20%): Scanning books for libraries, archival consulting

Governance:

  • Non-profit corporation (501(c)(3))
  • Board of directors (includes Brewster Kahle, founder)
  • Mission: “Universal access to all knowledge”

Sustainability:

  • Operating since 1996 (nearly 30 years)
  • Annual budget: ~$40 million
  • Endowment: Building toward long-term stability

Sovereignty Assessment:

  • Declaration: Free access, no user accounts required (can browse anonymously)
  • Connection: Open APIs, anyone can build on top of Archive data
  • Ground: Massive redundancy (multiple data centers, partner libraries), but centralized control

Challenges:

  • Legal vulnerability (2023: sued by publishers over book lending)
  • Funding concentration risk (what if major donors withdraw?)
  • Founder dependency (Brewster Kahle is central to organization)

Lesson: Non-profit model can sustain long-term preservation, but vulnerable to legal and funding risks.


Part III: Designing Sustainable Foundries §

The Foundry Business Canvas §

When designing a sovereignty business (we call these “Foundries”), use this canvas:

1. Value Proposition

  • What sovereignty problem do you solve?
  • For whom? (target users)
  • Why would they switch from incumbent platform?

2. Revenue Model

  • How do you make money?
  • Subscriptions? Hosting? Donations? Sales?
  • How much revenue per user? (unit economics)

3. Cost Structure

  • What are your main expenses?
  • Infrastructure (servers, bandwidth, storage)
  • Labor (developers, support, operations)
  • Legal/compliance

4. Three Pillars Integrity

  • Declaration: Do users own identities?
  • Connection: Can they communicate without surveillance?
  • Ground: Do they control their data?

5. Governance

  • Who makes decisions? (founders, board, users, workers?)
  • How democratic? (autocratic, representative, fully participatory?)
  • Exit strategy: What happens if founders leave/die?
  • For-profit, non-profit, co-op, hybrid?
  • What protects mission from capture?

7. Competitive Advantage

  • Why can’t megacorps copy you?
  • Is it technology, community, mission, legal structure?

8. Growth Strategy

  • How do you get first 100 users? 1,000? 10,000?
  • Network effects? (do you need them, or can you thrive small?)

9. Sustainability Timeline

  • Break-even: When do revenues exceed costs?
  • Maturity: When is the business self-sustaining?
  • Succession: How does it survive founders?

Example Canvas: Hypothetical “Sovereign Social” Platform §

1. Value Proposition:

  • Federated social network (like Mastodon) with managed hosting
  • Target: Non-technical users who want sovereignty but not self-hosting burden
  • Switch incentive: “Own your social media—no ads, no algorithm, no ban risk”

2. Revenue Model:

  • $10/month subscription per user
  • Includes: Custom domain (@[email protected]), 10GB storage, priority support

3. Cost Structure:

  • Infrastructure: $2/user/month (servers, bandwidth, storage)
  • Labor: $200k/year (2 developers, 1 support person)
  • Legal/admin: $20k/year
  • Break-even: 2,000 paying users ($240k/year revenue - $220k costs)

4. Three Pillars:

  • Declaration: Custom domains included, users control identity
  • Connection: ActivityPub federation (can follow/be followed from any compatible platform)
  • Ground: Full data export, can migrate to different host with all followers

5. Governance:

  • For-profit LLC initially (founders control)
  • Long-term: Convert to steward-ownership (founder gets salary, not equity) or co-op
  • User advisory board (elected representatives consult on policy)

6. Legal Structure:

  • Start: For-profit (easier to fund early development)
  • Mature: Steward-ownership (Purpose Foundation model) or B-Corp
  • Protection: Bylaws mandate Three Pillars compliance, can’t be removed

7. Competitive Advantage:

  • Can’t compete on features (Mastodon, Bluesky are free and feature-rich)
  • Advantage: Trust (users know they won’t be enshittified, locked in)
  • Niche: “We’re the managed Mastodon host for people who value sovereignty”

8. Growth Strategy:

  • Phase 1: 100 beta users (friends, early adopters) — $1k/month revenue
  • Phase 2: 1,000 users (content creators tired of platform instability) — $10k/month
  • Phase 3: 10,000 users (mainstream adoption) — $100k/month (profitable)
  • No VC needed (bootstrapped or small crowdfunding)

9. Sustainability Timeline:

  • Break-even: 2 years (2,000 users)
  • Maturity: 5 years (10,000 users, $1.2M/year revenue, stable team)
  • Succession: Founders create transition plan (documentation, steward-ownership transfer)

Viability Assessment:

  • Market: Small (most people don’t care about sovereignty)
  • But defensible: Those who do care are loyal, pay premium
  • Sustainable: Modest scale ($1-2M/year) is enough

Part IV: Case Studies in Foundry Economics §

Success: Basecamp (Bootstrapped, No VC) §

Business:

  • Project management software
  • Founded 1999 (as 37signals)
  • Never took VC funding

Revenue Model:

  • $99/month flat rate (unlimited users)
  • Simple pricing (no complex tiers)
  • Annual revenue: ~$50 million (estimated)

Economic Design:

  • Bootstrapped (profitable from early on)
  • Small team (~70 people) despite massive user base (thousands of companies)
  • No growth-at-all-costs (slow, steady, sustainable)

Sovereignty:

  • Declaration: Companies own their Basecamp data, own their URLs (custom domains available)
  • Connection: Not social, so less relevant (but API for integrations)
  • Ground: Good data export, can migrate to self-hosted alternatives if needed

Key Lesson: Profitability at small scale (relative to VC-funded competitors) enables sovereignty.

Why It Works:

  • Founders (Jason Fried, DHH) ideologically opposed to VC
  • Company structure allows them to say no to growth pressure
  • Loyal user base willing to pay premium for stability

Failure: Ello (VC-Funded “Anti-Facebook”) §

Premise (2014):

  • Social network promising “no ads, no data mining”
  • Launched as alternative to Facebook
  • Tagline: “You are not a product”

Business Model:

  • Initially free
  • Plan: Freemium (paid features: analytics, custom domains, themes)

Funding:

  • Raised $5.5 million VC funding (Series A)
  • Converted to Public Benefit Corporation (B-Corp) to enshrine ad-free mission

What Went Wrong:

  • VC pressure to grow fast (needed massive user base to justify valuation)
  • Network effects didn’t materialize (no one on Ello = no reason to join Ello)
  • Pivot to niche (2016): Became platform for artists/creators only
  • Acquired 2021 by holding company; original mission abandoned

Sovereignty Failure:

  • Despite B-Corp status, VC funding created growth pressure
  • Users who joined believing in mission felt betrayed by pivot
  • Platform never achieved critical mass for sustainability

Key Lesson: VC funding is incompatible with sovereignty, even with legal protections.

Partial Success: Mastodon (Donations + Volunteer Labor) §

Business Model:

  • Mastodon is open-source software (free)
  • Creator (Eugen Rochko) funded by Patreon donations
  • Instance hosting is decentralized (thousands of admins, each with own funding model)

Rochko’s Income:

  • Patreon: $30k/month from ~6,000 patrons
  • Grants: Occasional from Mozilla, NGI
  • Annual: ~$400k (modest for software developer in US, but sustainable)

Instance Funding (Varied):

  • Some free (admin pays out of pocket)
  • Some donation-supported (Patreon, Ko-fi)
  • Some subscription ($5-10/month per user)
  • Some institutionally backed (universities, non-profits)

Sustainability Assessment:

  • Core software: Sustainable (Rochko funded, plus volunteer contributors)
  • Instances: Fragile (many admins burn out, shut down)
  • Overall: Network survives because federated (if one instance dies, users migrate)

Sovereignty:

  • Excellent (federated, open protocol, self-hostable)

Key Lesson: Donation-funded + decentralized can work, but creates admin burnout risk.


Part V: Avoiding the Failure Modes §

Failure Mode 1: The Heroic Founder Problem §

Symptom:

  • Organization depends on one person (founder/maintainer)
  • If they burn out, die, or leave, project collapses

Examples:

  • Small open-source projects (single maintainer)
  • Volunteer-run archives (when admin quits, archive vanishes)

Prevention:

  • Build team, not solo operation
  • Document everything (so others can take over)
  • Succession planning (who’s next in charge?)
  • Institutional structure (legal entity that outlives founder)

Failure Mode 2: The Volunteer Burnout Trap §

Symptom:

  • Project relies on unpaid labor
  • Initial enthusiasm fades
  • No one has time/energy to maintain

Examples:

  • Mastodon instances (many shut down after 1-2 years)
  • Open-source projects (maintainers quit from exhaustion)

Prevention:

  • Pay people (even modest stipends help)
  • Limit scope (don’t promise more than you can sustain)
  • Rotate responsibilities (avoid single points of failure)

Failure Mode 3: Speculative Capture §

Symptom:

  • Company/protocol gets bought by VCs, megacorps, or speculators
  • New owners prioritize profit over mission
  • Enshittification follows

Examples:

  • Instagram (acquired by Facebook)
  • Tumblr (Yahoo, then Verizon, then Automattic)
  • Many blockchain projects (early idealism → speculative frenzy)

Prevention:

  • Legal structure that prevents sale (non-profit, co-op, steward-ownership)
  • Open source (so community can fork if captured)
  • Mission codification (bylaws that can’t be changed)

Failure Mode 4: Complexity Collapse §

Symptom:

  • System becomes too complex to maintain
  • Technical debt accumulates
  • Eventually, no one understands how it works

Examples:

  • Legacy software (COBOL banking systems)
  • Over-engineered platforms (added features until bloated)

Prevention:

  • Simplicity as core value (resist feature creep)
  • Regular refactoring (pay down technical debt)
  • Documentation (explain how things work)

Part VI: The 10-Year Business Plan §

If you’re building a Foundry, plan for the long haul:

Year 1: Proof of Concept §

  • Build MVP (minimum viable product)
  • Get 10-50 early users
  • Validate that people will pay
  • Revenue: $0-5k/year

Years 2-3: Find Product-Market Fit §

  • Iterate based on user feedback
  • Grow to 100-500 users
  • Break even or close to it
  • Revenue: $10k-50k/year

Years 4-5: Scale to Sustainability §

  • 1,000-5,000 users
  • Profitable (revenue exceeds costs)
  • Hire small team (2-5 people)
  • Revenue: $100k-500k/year

Years 6-10: Mature and Institutionalize §

  • 5,000-50,000 users
  • Diversify revenue (not dependent on single income stream)
  • Succession planning (ensure survival past founders)
  • Revenue: $500k-5M/year

Beyond Year 10: Legacy §

  • Convert to permanent structure (co-op, foundation, steward-ownership)
  • Ensure mission persists even if founders leave
  • Document everything for future maintainers

Key Insight: You don’t need billions of users or unicorn valuation. Small, sustainable, and sovereign is success.


Conclusion: The Anvil That Endures §

Building a Foundry is hard. You’re competing against platforms with billions in VC funding, network effects, and zero regard for user sovereignty.

But you have advantages they don’t:

  • Trust: Users know you won’t betray them
  • Longevity: You’re building for 50 years, not next quarter
  • Mission: You care about sovereignty, not extraction

The economics of the Anvil require patience:

  • Growth will be slow
  • You’ll never be “unicorn” rich
  • You’ll always be outspent by megacorps

But you’ll build something that lasts. Something that users own. Something that can’t be murdered by a quarterly earnings call.

The Anvil endures not because it grows fastest, but because it’s built to survive.

In the next chapter, we explore distributed commons governance—how to build infrastructure that many organizations share, using Elinor Ostrom’s principles for managing common-pool resources.

For now, sketch your Foundry. What would you build? How would you fund it? And how would you ensure it embodies the Three Pillars while remaining economically viable?

The Anvil awaits the forging.


Discussion Questions §

  1. VC Dilemma: If you had a great idea for a sovereign platform but needed $1M to build it, would you take VC funding? Why or why not? What alternatives exist?

  2. Subscription Exclusion: User-pays models exclude people who can’t afford subscriptions. Is this an acceptable trade-off for sovereignty? How could you address it?

  3. Co-op Governance: Would you want to run a platform democratically (co-op model)? What are the benefits and frustrations of democratic governance?

  4. Small vs. Big: Is it better to be small and sovereign (10,000 loyal users) or big and compromised (100 million users but VC-funded)? Does scale matter?

  5. Competition: How do you compete with “free” platforms (Gmail, Facebook, Instagram) when you charge money? What’s your value proposition?

  6. Your Own Business: If you built a Foundry, what would your revenue model be? Walk through the Business Canvas for your hypothetical platform.


Exercise: Design Your Foundry §

Task: Design a complete business plan for a sovereignty-respecting platform.

Part 1: The Problem (300 words)

  • What platform are you replacing/competing with?
  • What sovereignty violations does it commit?
  • Who’s your target user? (specific niche, not “everyone”)

Part 2: The Business Canvas (1000 words)

Complete all 9 sections:

  1. Value Proposition
  2. Revenue Model (with unit economics)
  3. Cost Structure
  4. Three Pillars Integrity Check
  5. Governance Model
  6. Legal Structure
  7. Competitive Advantage
  8. Growth Strategy (with realistic numbers)
  9. Sustainability Timeline

Part 3: 5-Year Financial Projection (500 words)

Create a simple table: | Year | Users | Revenue/User | Total Revenue | Total Costs | Profit/Loss | |------|-------|--------------|---------------|-------------|-------------| | 1 | 50 | $120/year | $6k | $50k | -$44k | | 2 | 500 | $120/year | $60k | $100k | -$40k | | 3 | 2,000 | $120/year | $240k | $200k | +$40k | | 4 | 5,000 | $120/year | $600k | $400k | +$200k | | 5 | 10,000| $120/year | $1.2M | $700k | +$500k |

Explain assumptions. When do you break even? Is this realistic?

Part 4: Failure Mode Analysis (500 words)

  • What’s your biggest risk? (heroic founder, burnout, capture, complexity?)
  • How do you mitigate it?
  • What’s your “if this fails” exit strategy? (can users take their data elsewhere?)

Part 5: Reflection (300 words)

  • Would you actually want to run this business?
  • What’s the hardest part?
  • What did you learn about the tensions between sovereignty and economics?

Further Reading §

On Platform Economics §

  • Doctorow, Cory. “Competitive Compatibility: Let’s Fix the Internet, Not the Tech Giants.” Electronic Frontier Foundation (2019).
  • Srnicek, Nick. Platform Capitalism. Polity, 2017.
  • Zuboff, Shoshana. The Age of Surveillance Capitalism. PublicAffairs, 2019.

On Alternative Business Models §

  • Schneider, Nathan. “An Internet of Ownership.” Sociological Review 68, no. 2 (2020).
  • Scholz, Trebor. Platform Cooperativism. Rosa Luxemburg Stiftung, 2016.
  • Muldoon, James. Platform Socialism. Pluto Press, 2022.

On Sustainable Funding §

  • Eghbal, Nadia. Working in Public: The Making and Maintenance of Open Source Software. Stripe Press, 2020.
  • On how open source projects fund themselves

  • Benkler, Yochai. The Wealth of Networks. Yale University Press, 2006.

  • On peer production and non-market economics

On Company Case Studies §

  • Fried, Jason, and DHH. Rework. Crown Business, 2010.
  • Basecamp’s philosophy (bootstrapped, profitable, sovereign)

  • Rochko, Eugen. “Mastodon Blog.” https://blog.joinmastodon.org

  • Founder’s posts on building/funding decentralized platform

On Business Design §

  • Osterwalder, Alexander, and Yves Pigneur. Business Model Generation. Wiley, 2010.
  • Business canvas methodology

  • Purpose Foundation. “Steward-Ownership.” https://purpose-economy.org/en/

  • Alternative ownership structures

End of Chapter 12

Next: Chapter 13 — Distributed Commons Governance (The Seed Bank)

Part III • Institutional Design, Commons & Political Economy

Chapter 13: Distributed Commons Governance

Building the Seed Bank

26 min read 5,657 words

Opening: The Problem of Scale §

In 2014, the Internet Archive held approximately 15 petabytes of data—one of the largest digital collections in the world. Impressive. Essential. But also: terrifying.

All that cultural memory, concentrated in one organization, in two physical locations (San Francisco and Alexandria). What if:

  • A fire destroys the data centers?
  • A lawsuit bankrupts the organization?
  • A government decides to shut it down?
  • Climate change floods the facilities?
  • A cyberattack encrypts everything?

Brewster Kahle, founder of the Internet Archive, knows this risk. He’s said publicly: “We need more Internet Archives.” Not mirrors of the Internet Archive—but independent preservation organizations running parallel efforts, creating redundancy.

But here’s the problem: preservation at scale requires collective action. One person can’t preserve the internet. One organization can struggle but ultimately faces existential risks. We need many organizations working together—a distributed commons for digital preservation.

Yet commons are famously unstable. Garrett Hardin’s “Tragedy of the Commons” (1968) argues that shared resources inevitably collapse: everyone takes, no one maintains, the commons degrades until it’s useless.

This chapter asks: How do we build a Seed Bank—a distributed network of preservation nodes that cooperate to preserve digital culture—without falling into tragedy of the commons?

The answer comes from Elinor Ostrom, who proved Hardin wrong. Commons can be governed sustainably—if designed correctly. This chapter applies Ostrom’s principles to digital preservation, showing how to build governance systems that resist collapse.

We’ll explore:

  • Ostrom’s 8 Design Principles for sustainable commons
  • Case studies: LOCKSS (success), Mastodon (challenges), Software Heritage (academic model)
  • How to design governance for distributed preservation
  • Technical architecture for Seed Banks (distributed storage, federated governance)
  • Why this is harder than it sounds (and how to succeed anyway)

Part I: Elinor Ostrom and the Governing the Commons §

The Tragedy of the Commons (Wrong) §

Garrett Hardin’s Argument (1968):

  • Imagine a pasture shared by herders
  • Each herder benefits from adding another cow (more milk/meat)
  • Cost of overgrazing is shared among all herders
  • Rational self-interest → everyone adds cows → pasture collapses
  • Solution: Private property or government control

Applied to Digital Preservation:

  • Internet Archive preserves web (public good)
  • Everyone benefits from using it
  • No one pays (free access)
  • Costs are borne by one organization
  • Eventually: Funding crisis, collapse

Hardin’s logic suggests: Digital preservation commons can’t work. Either privatize it (paywalls, licenses) or nationalize it (government mandate, taxes).

Ostrom’s Rebuttal (Right) §

Elinor Ostrom’s Research (1990):

  • Studied commons that didn’t collapse: irrigation systems in Spain, forests in Japan, fisheries in Maine
  • Found: Commons governed by communities (not private owners or states) can be sustainable
  • Key: Design principles that prevent overuse and ensure maintenance

Won Nobel Prize in Economics (2009) for proving Hardin wrong.

Applied to Digital Preservation:

  • Distributed preservation networks (Seed Banks) can work
  • If designed with Ostrom’s principles
  • Communities of preservation organizations can self-govern
  • No need for monopoly (Internet Archive) or government takeover

Ostrom’s 8 Design Principles §

Ostrom identified eight characteristics of sustainable commons:

  1. Clearly Defined Boundaries
  2. Proportionality Between Benefits and Costs
  3. Collective Choice Arrangements
  4. Monitoring
  5. Graduated Sanctions
  6. Conflict Resolution Mechanisms
  7. Minimal Recognition of Rights
  8. Nested Enterprises (for large-scale commons)

Let’s examine each principle and how it applies to digital preservation.


Part II: Applying Ostrom’s Principles to Digital Preservation §

Principle 1: Clearly Defined Boundaries §

Ostrom’s Principle:

  • Who has rights to use the commons? (Clear membership)
  • What are the boundaries of the resource? (What’s included/excluded?)

Why It Matters:

  • Without boundaries, outsiders can exploit the commons without contributing
  • Without clear resource definition, disputes arise over what’s being governed

Applied to Seed Bank:

Who can participate?

  • Define membership criteria
  • Universities with archival capacity?
  • Non-profits with preservation mandates?
  • Volunteers meeting technical requirements?
  • Not open to anyone (risk of bad actors overwhelming system)
  • But not so exclusive that you can’t scale

What’s being preserved?

  • Scope of the commons: All web content? Specific platforms? Specific geographies?
  • Clear policy on what gets preserved (use Custodial Filter from Chapter 5)
  • Boundaries prevent mission creep (“We preserve murdered platforms, not general web archiving”)

Example: LOCKSS (Lots of Copies Keep Stuff Safe)

Membership:

  • University libraries and academic institutions
  • Must meet technical requirements (storage, bandwidth, uptime)
  • Must commit to preservation mandate (not just using it for free)

Resource Boundaries:

  • Academic journals and books
  • Not general web content (that’s Internet Archive’s domain)
  • Clear scope prevents overlap and confusion

Result: LOCKSS has run for 20+ years with 300+ participating libraries. Boundaries work.

Principle 2: Proportionality Between Benefits and Costs §

Ostrom’s Principle:

  • Costs of maintaining the commons should be proportional to benefits received
  • Those who use more should contribute more

Why It Matters:

  • If heavy users don’t pay their share, resentment builds, cooperation collapses
  • Freeloaders undermine collective will to maintain commons

Applied to Seed Bank:

Challenge: Digital preservation has unusual economics:

  • Marginal cost of one more user accessing data is near-zero (unlike grazing land, which depletes)
  • But infrastructure costs are real: storage, bandwidth, maintenance

Proportionality Mechanisms:

1. Storage-Based Contributions

  • Organizations contribute storage proportional to what they preserve
  • If you preserve 1TB, you provide 1TB+ of storage (redundancy)

2. Bandwidth-Based Contributions

  • Heavy downloaders provide bandwidth to others (BitTorrent model)
  • Upload/download ratios

3. Labor Contributions

  • Some orgs provide storage, others provide metadata curation, others provide technical development
  • Value different contributions (not just storage)

4. Financial Sliding Scale

  • Wealthy universities pay more, small non-profits pay less
  • But everyone contributes something (even if token)

Example: LOCKSS Implementation

Costs:

  • Libraries pay annual membership fee (sliding scale based on size)
  • Provide servers and storage (technical contribution)
  • Participate in governance (labor contribution)

Benefits:

  • Access to entire LOCKSS archive
  • Redundancy for their own collections (others preserve their journals)
  • Collective preservation cheaper than individual efforts

Proportionality: Large research universities pay more, small colleges pay less, but all contribute. Balanced.

Principle 3: Collective Choice Arrangements §

Ostrom’s Principle:

  • People affected by rules should participate in making/modifying them
  • Not top-down imposition—democratic or consensus-based governance

Why It Matters:

  • Rules imposed externally are resented and resisted
  • Collective choice creates buy-in and legitimacy

Applied to Seed Bank:

Who decides:

  • What gets preserved? (Content policy)
  • How it’s preserved? (Technical standards)
  • Who gets access? (Public, researchers, restricted?)
  • How to allocate resources? (Storage priorities)

Governance Models:

1. One Member, One Vote

  • All participating organizations have equal say
  • Democratic but slow
  • Risk: Large and small orgs weighted equally (is that fair?)

2. Weighted Voting

  • Vote weight proportional to contribution (storage, funding, labor)
  • More equitable but complex
  • Risk: Wealthy orgs dominate

3. Consensus Decision-Making

  • Major decisions require consensus (not just majority)
  • Ensures minority voices heard
  • Risk: Gridlock

4. Federated Councils

  • Representatives from subgroups (geographic regions, institution types, technical roles)
  • Balances representation with efficiency

Example: Mastodon’s Governance Struggles

Mastodon Approach:

  • Each instance governed by its admin(s)
  • No central governance over whole network
  • ActivityPub protocol decisions made by W3C (external standards body)

Problems:

  • No mechanism for collective decision on network-wide issues (moderation, defederation policies)
  • Admins burn out (all burden on individuals)
  • Large instances dominate (network effects concentrate power despite federation)

Lesson: Collective choice requires structure. Pure decentralization without governance mechanisms fails.

Principle 4: Monitoring §

Ostrom’s Principle:

  • Someone must monitor compliance with rules
  • Monitors should be accountable to users (not external authorities)

Why It Matters:

  • Without monitoring, freeloaders go undetected
  • But monitoring by external authorities (police, government) breeds resentment

Applied to Seed Bank:

What to Monitor:

1. Technical Compliance

  • Are nodes providing promised storage?
  • Is data being preserved correctly (checksums, bit rot detection)?
  • Are nodes online and accessible?

2. Participation

  • Are members contributing labor (metadata, curation)?
  • Are they participating in governance (voting, meetings)?

3. Ethical Compliance

  • Are members following Custodial Filter? (Not preserving harmful content in violation of policy)
  • Respecting access restrictions?

How to Monitor:

Automated Technical Monitoring:

  • Software checks: Are nodes responding? Are checksums valid?
  • Bandwidth and storage metrics
  • Alerts when nodes fail

Peer Monitoring:

  • Members audit each other (rotating responsibility)
  • Transparent metrics (everyone can see who’s contributing)
  • Community accountability (not centralized enforcement)

Example: LOCKSS’s Polling System

Mechanism:

  • LOCKSS nodes periodically “poll” each other: “Do you have this file? Is your checksum correct?”
  • If discrepancies detected, nodes vote: Which version is correct?
  • Majority consensus repairs corrupted copies
  • No central authority—peer-to-peer verification

Result: Automated monitoring + collective verification. No single point of failure.

Principle 5: Graduated Sanctions §

Ostrom’s Principle:

  • Rule violations should be met with sanctions
  • Sanctions should escalate: warning → fine → suspension → expulsion
  • Not immediate harsh punishment (which breeds resentment)

Why It Matters:

  • Without sanctions, rules are meaningless
  • But overly harsh sanctions create fear and reduce cooperation

Applied to Seed Bank:

Violation Types:

Minor Violations:

  • Missing a governance meeting
  • Temporary technical downtime (server maintenance)
  • Late financial contribution

Moderate Violations:

  • Persistent non-participation
  • Repeated technical failures (unreliable node)
  • Minor policy violations (preserving out-of-scope content)

Major Violations:

  • Deliberately preserving harmful content in violation of Custodial Filter
  • Attempting to monetize shared data (violating commons ethos)
  • Sabotage (deleting others’ data)

Graduated Response:

1st Violation (Minor): Warning, offer support (maybe they need technical help)

2nd Violation (Moderate): Formal reprimand, reduce privileges (e.g., lower storage quota)

3rd Violation (Major): Suspension (temporary loss of access and voting rights)

4th Violation (Severe): Expulsion (removed from network)

Appeal Process:

  • Members can appeal sanctions
  • Neutral arbitration panel reviews

Example: Academic Consortium Models

Many academic consortia (library networks, research cooperatives) use graduated sanctions:

  • First violation → email reminder
  • Second → formal letter from consortium director
  • Third → loss of specific benefits (can’t borrow from other libraries)
  • Fourth → expulsion (rare, reserved for egregious violations)

Key: Sanctions are restorative, not purely punitive. Goal is to bring violators back into compliance, not to purge members.

Principle 6: Conflict Resolution Mechanisms §

Ostrom’s Principle:

  • Disputes will arise—need fast, low-cost, legitimate ways to resolve them
  • Local resolution better than external courts

Why It Matters:

  • Without conflict resolution, disputes fester, cooperation breaks down
  • Expensive litigation destroys commons (costs exceed benefits)

Applied to Seed Bank:

Common Disputes:

1. Resource Allocation

  • “Why does University X get more storage than us?”
  • “Our node is down; who’s responsible for lost data?”

2. Content Disputes

  • “Should we preserve this controversial content?”
  • “Someone violated the Custodial Filter; what do we do?”

3. Governance Disputes

  • “This policy was passed unfairly; some members weren’t consulted”
  • “The voting process is biased toward large institutions”

4. Technical Disputes

  • “Node Y isn’t maintaining their checksums correctly”
  • “Our data was corrupted; who’s liable?”

Resolution Mechanisms:

Tier 1: Direct Negotiation

  • Parties try to resolve dispute themselves
  • Encouraged before escalation

Tier 2: Mediation

  • Neutral member mediates
  • Non-binding (parties can reject mediation outcome)

Tier 3: Arbitration

  • Panel of 3-5 members hears case
  • Binding decision (parties agree to abide by outcome)
  • Faster and cheaper than courts

Tier 4: External Courts (last resort)

  • Only for major legal issues (breach of contract, fraud)
  • Avoided if possible (expensive, slow, undermines commons)

Example: Wikipedia’s Dispute Resolution

Wikipedia has multi-tiered conflict resolution:

  • Direct talk page discussion
  • Third Opinion (neutral editor weighs in)
  • Requests for Comment (community input)
  • Arbitration Committee (binding decision)

Lesson: Most disputes resolved at lower tiers. Arbitration is rare. System works because it’s fast, low-cost, and legitimate (community-run, not external).

Principle 7: Minimal Recognition of Rights §

Ostrom’s Principle:

  • External authorities (government, courts) should recognize the community’s right to self-govern
  • Don’t need full legal sovereignty, but need enough autonomy to enforce rules

Why It Matters:

  • If external authorities constantly override community rules, self-governance is impossible
  • Need legal protection from outsiders who want to undermine or destroy the commons

Applied to Seed Bank:

What Rights Are Needed?

1. Right to Exist

  • Legal recognition as an entity (non-profit, cooperative, consortium)
  • Can enter contracts, own property (servers, storage)

2. Right to Make Rules

  • Can set membership criteria, content policies, technical standards
  • Not overridden by government unless violating law

3. Right to Exclude

  • Can remove bad actors
  • Not forced to include everyone (boundaries matter)

4. Right to Fair Use / Preservation

  • Legal protection for preservation activities (scraping, format migration)
  • Copyright exceptions for archival purposes

5. Right to Federate

  • Can form partnerships with other preservation networks
  • Not locked into national boundaries or single legal jurisdiction

Threats to Rights:

Legal:

  • Copyright lawsuits (preserving content without permission)
  • DMCA takedowns (if content violates IP law)
  • Platform Terms of Service (scraping prohibited, legal gray area)

Political:

  • Government censorship (forced to remove content)
  • National security claims (data seizure)
  • Taxation or regulation that makes preservation unaffordable

How to Secure Rights:

1. Legal Structuring

  • Incorporate as 501(c)(3) non-profit (US) or charitable trust (UK) or equivalent
  • Protects from certain liabilities, provides tax advantages

2. Advocacy

  • Lobby for “Right to Archive” laws
  • Expand fair use for preservation
  • Fight restrictive copyright (Section 1201 of DMCA in US)

3. International Cooperation

  • Distribute nodes globally (no single government can shut down entire network)
  • Partner with organizations in multiple jurisdictions

Example: Internet Archive’s Legal Battles

Challenges:

  • Sued by publishers over Controlled Digital Lending (book scanning)
  • Frequent DMCA takedowns for archived web content
  • Threatened by record labels over audio preservation

Defense:

  • Fair use arguments (preservation is non-commercial, transformative)
  • Public advocacy (builds political support)
  • International presence (if US law becomes hostile, shift focus elsewhere)

Lesson: Commons need legal protection. Pure grassroots self-governance isn’t enough if external authorities can destroy you.

Principle 8: Nested Enterprises (For Large Commons) §

Ostrom’s Principle:

  • For large commons, organize in nested layers
  • Local decisions at local level, regional at regional, global at global
  • Subsidiarity: Decisions made at lowest effective level

Why It Matters:

  • Single governance structure for massive commons doesn’t scale
  • Nested layers allow local autonomy while maintaining coordination

Applied to Seed Bank:

Nested Structure Example:

Layer 1: Individual Nodes

  • Single institution (university, non-profit, library)
  • Runs own servers, decides local policies (what to prioritize, how much storage to allocate)
  • Autonomous within consortium guidelines

Layer 2: Regional Consortia

  • Groups of nodes in same geography (e.g., “Northeast US Consortium,” “European Seed Bank Alliance”)
  • Coordinate regional priorities, share resources, handle regional disputes

Layer 3: Global Network

  • All regional consortia coordinate
  • Set global standards (technical protocols, ethical guidelines)
  • Handle cross-regional issues (international copyright, data sovereignty)

Decision Allocation:

Local Level:

  • Which specific content to preserve
  • Technical implementation details (hardware, software)
  • Day-to-day operations

Regional Level:

  • Resource sharing within region
  • Regional content priorities (e.g., European consortium prioritizes European platforms)
  • Regional legal compliance

Global Level:

  • Network-wide standards and protocols
  • Conflict resolution between regions
  • Major policy changes (ethics, access, membership criteria)

Example: LOCKSS’s Nested Structure

Individual Libraries:

  • Run LOCKSS boxes (servers)
  • Decide what to preserve (within LOCKSS framework)

LOCKSS Alliance:

  • Consortium of participating libraries
  • Coordinates technical standards, shares metadata

LOCKSS Program (at Stanford):

  • Central organization providing software and coordination
  • Not top-down control—facilitative role

Result: Local autonomy + global coordination. Libraries aren’t dictated to, but also aren’t isolated.


Part III: Case Studies in Distributed Commons §

Case Study 1: LOCKSS (Success Story) §

LOCKSS = Lots of Copies Keep Stuff Safe

Launched: 1999 (Stanford University Libraries)

Mission: Distributed digital preservation for academic journals and books

Ostrom Principles Implementation:

1. Boundaries:

  • Membership: Academic libraries meeting technical criteria
  • Resource: Scholarly publications (not general web)

2. Proportionality:

  • Libraries contribute storage proportional to usage
  • Sliding scale membership fees

3. Collective Choice:

  • Governance board includes library representatives
  • Major decisions voted on by members

4. Monitoring:

  • Automated polling system (nodes verify each other’s data)
  • Transparent metrics (uptime, storage, participation)

5. Graduated Sanctions:

  • Non-compliant nodes warned, then suspended, then expelled (rare)

6. Conflict Resolution:

  • Disputes handled by governance board
  • Mediation before arbitration

7. Minimal Recognition:

  • Non-profit consortium legally recognized
  • Fair use protections for preservation

8. Nested Enterprises:

  • Individual libraries → Regional networks → Global LOCKSS Alliance

Results:

  • 300+ participating libraries worldwide
  • 20+ years of stable operation
  • Petabytes of preserved content
  • No major tragedies of commons

Why It Worked:

  • Clear mission and boundaries
  • Strong technical foundation (automated monitoring, peer verification)
  • Academic culture of cooperation (libraries already collaborate)
  • Sustainable funding (membership fees + grants)

Lessons: Ostrom’s principles work. But require:

  • Careful design from start
  • Ongoing maintenance of governance
  • Cultural fit (participants value commons)

Case Study 2: Mastodon (Mixed Results) §

Mastodon: Federated social network (launched 2016)

Model: Anyone can run an instance; instances federate via ActivityPub

Ostrom Principles Analysis:

1. Boundaries:Weak

  • Anyone can start an instance (no membership criteria)
  • No clear definition of “what is Mastodon network” (any ActivityPub server can join)
  • Result: Toxic instances proliferate, defederation wars

2. Proportionality:Absent

  • Users on large instances consume resources (bandwidth, moderation) but don’t contribute
  • Small instances subsidize large ones (infrastructure costs borne unevenly)
  • No mechanism to enforce proportional contribution

3. Collective Choice: ⚠️ Fragmented

  • Each instance admin makes rules for their instance
  • No network-wide governance (deliberate choice, but creates problems)
  • Major decisions (protocol changes) made by W3C (external body)

4. Monitoring: ⚠️ Limited

  • No network-wide monitoring (each instance monitors itself)
  • Bad actors can spin up new instances faster than they’re defederated

5. Graduated Sanctions:Absent

  • Only tool: Defederation (nuclear option—completely sever connection)
  • No middle ground (warning, temporary suspension, etc.)

6. Conflict Resolution:Absent

  • No formal mechanism for resolving disputes between instances
  • Admins handle conflicts ad hoc (often poorly)

7. Minimal Recognition:Strong

  • Legally: Mastodon is just software; instances are independent entities
  • No central organization to sue or shut down

8. Nested Enterprises:Absent

  • Flat structure (instances federate directly, no regional coordination)
  • No higher-level governance

Results:

  • Rapid growth (millions of users)
  • But: Admin burnout, moderation nightmares, defederation drama
  • Large instances recentralize (Mastodon.social dominates)
  • Network effects undermine federation (most users on a few big instances)

Why It Struggled:

  • No governance designed in—assumed federation = automatic self-governance
  • Ostrom’s principles ignored (implicitly or explicitly)
  • Result: Some of Hardin’s predictions came true (overuse, collapse of cooperation)

Lessons:

  • Pure decentralization without governance doesn’t work
  • Need explicit commons governance, not assumption that “protocol solves it”
  • Federation is necessary but not sufficient

Case Study 3: Software Heritage (Academic Model) §

Software Heritage: Preserving all open-source software (launched 2016)

Model: Academic consortium funded by Inria (French research institute) + partners

Ostrom Principles Implementation:

1. Boundaries:Clear

  • Resource: Open-source software (not proprietary)
  • Membership: Academic institutions and non-profits committed to preservation

2. Proportionality: ⚠️ Weak

  • Inria provides most funding (imbalanced)
  • Contributors provide mirrors (storage) but not all equally
  • Needs better proportionality as it scales

3. Collective Choice: ⚠️ Limited

  • Governance by Inria + advisory board
  • Not fully democratic (participants have input but Inria has veto)
  • Transitioning to more participatory model

4. Monitoring:Strong

  • Automated crawlers monitor GitHub, GitLab, etc.
  • Checksums verify integrity
  • Mirrors regularly audited

5. Graduated Sanctions:N/A (so far)

  • No major violations yet (early stage)

6. Conflict Resolution: ⚠️ Informal

  • Academic disputes handled through traditional academic channels
  • No formal arbitration process

7. Minimal Recognition:Strong

  • UNESCO partnership (international recognition)
  • French government support (legal standing)

8. Nested Enterprises: ⚠️ Emerging

  • Inria (central) → Partner institutions (regional) → Mirrors (local)
  • Structure is forming but not fully nested yet

Results:

  • 15+ billion source code files archived
  • Growing academic and industry support
  • Stable funding (for now—depends on Inria)

Why It Works (So Far):

  • Strong institutional backing (Inria)
  • Clear mission and technical competence
  • Academic culture of openness

Vulnerabilities:

  • Funding concentration (too dependent on Inria)
  • Governance not fully participatory (top-down elements)
  • Needs to scale proportionality and nested governance

Lessons:

  • Academic commons can work but need explicit governance
  • Early centralization (Inria) was pragmatic (fast start) but must transition to distributed governance for long-term sustainability

Part IV: Technical Architecture for the Seed Bank §

Distributed Storage Models §

Challenge: How do you technically implement a Seed Bank where many organizations cooperate?

Models:

1. Peer-to-Peer (BitTorrent-style)

How It Works:

  • Each node stores chunks of data
  • Nodes share chunks with each other
  • Redundancy through replication (each chunk stored on multiple nodes)

Pros:

  • Highly resilient (no single point of failure)
  • Scales with number of nodes (more nodes = more capacity)

Cons:

  • Coordination complexity (how to ensure chunks are distributed evenly?)
  • Discovery problem (how do you find what you need?)
  • Freeloaders (leechers who download but don’t upload)

Example: IPFS (InterPlanetary File System)

  • Content-addressed storage (files identified by hash, not location)
  • Nodes pin content they care about
  • Network collectively preserves pinned content

For Seed Bank:

  • Works well for technical infrastructure
  • But needs governance layer (Ostrom principles) to prevent freeloading

2. Federated Repositories

How It Works:

  • Each organization runs a full repository (or mirrors subset)
  • Repositories sync with each other periodically
  • Metadata standardized (everyone knows what everyone else has)

Pros:

  • Each node is autonomous (can operate independently)
  • Clear responsibility (each org manages its own repository)

Cons:

  • Requires significant resources per node (each must store large amounts)
  • Sync complexity (keeping repositories aligned)

Example: LOCKSS

  • Each library runs a LOCKSS box
  • Boxes poll each other to verify integrity
  • If one box fails, others have copies

For Seed Bank:

  • Best model for institutional commons
  • Matches Ostrom’s principles (clear boundaries, monitoring, etc.)

3. Centralized Coordination, Distributed Storage

How It Works:

  • Central registry tracks what each node stores (metadata)
  • Nodes store actual data (distributed)
  • Central registry doesn’t store content (only pointers)

Pros:

  • Easy discovery (query central registry: “who has this file?”)
  • Nodes can specialize (some store rare items, others common items)

Cons:

  • Central registry is single point of failure (for discovery, not storage)
  • Risk of centralization creep (registry gains too much power)

Example: Archive.org + mirrors

  • Internet Archive is primary (central)
  • Mirrors exist globally (distributed storage)
  • Archive.org coordinates but doesn’t control mirrors

For Seed Bank:

  • Pragmatic hybrid
  • But must ensure central registry doesn’t become bottleneck or dictator

Governance-Integrated Architecture §

Key Insight: Governance can’t be afterthought—must be built into technical architecture.

Design Patterns:

Smart Contracts for Proportionality

Use blockchain/smart contracts to enforce:

  • Storage contributions (can’t withdraw more than you deposit)
  • Bandwidth limits (upload/download ratios)
  • Automated sanctions (node falls below threshold → reduced access)

Pros: Self-enforcing, transparent, automated

Cons: Requires crypto infrastructure (complexity, cost), immutable (hard to change rules)

Best for: Technical compliance (storage, bandwidth, uptime)

Voting Mechanisms in Protocol

Embed governance in protocol:

  • Major changes require majority vote of nodes
  • Nodes vote by running updated software
  • Forks if consensus fails (like blockchain forks)

Pros: Democratic, decentralized

Cons: Slow (consensus takes time), can fork network

Best for: Major protocol changes (technical standards, access policies)

Reputation Systems

Track node behavior:

  • Nodes earn reputation for uptime, correct checksums, participation
  • High-reputation nodes get priority (bandwidth, storage)
  • Low-reputation nodes sanctioned (reduced access, eventually expelled)

Pros: Incentivizes good behavior, graduated sanctions

Cons: Gameable (Sybil attacks, reputation washing), requires trusted reputation oracle

Best for: Monitoring and sanctions (Ostrom principles 4-5)


Part V: Building Your Own Seed Bank §

Step-by-Step: Launching a Distributed Preservation Network §

Phase 1: Coalition Formation (Year 1)

Goals:

  • Recruit 5-10 founding members (organizations committed to preservation)
  • Establish shared mission and values
  • Draft initial governance charter

Activities:

  1. Identify potential partners:
  2. Universities with archival programs
  3. Libraries with digital collections
  4. Non-profits focused on preservation (regional Internet Archives, etc.)

  5. Host founding workshop:

  6. Day-long meeting to discuss vision, values, governance
  7. Agree on Ostrom principles implementation
  8. Sign founding charter

  9. Secure seed funding:

  10. Apply for grants (Mellon, Knight, NEH)
  11. Pitch: “Building distributed commons for digital preservation”
  12. Initial funding for technical infrastructure + coordination

Phase 2: Technical Infrastructure (Year 1-2)

Goals:

  • Deploy technical infrastructure (storage, networking, monitoring)
  • Pilot with small collection (test the system)

Activities:

  1. Choose technical model:
  2. Federated repositories (LOCKSS-style)?
  3. P2P (IPFS-style)?
  4. Hybrid (centralized coordination, distributed storage)?

  5. Deploy pilot:

  6. Each founding member sets up node
  7. Preserve test collection (e.g., one murdered platform’s archive)
  8. Verify redundancy, checksums, access

  9. Build monitoring systems:

  10. Automated health checks (nodes online?)
  11. Peer verification (checksums correct?)
  12. Dashboard showing network status

Phase 3: Governance Formalization (Year 2)

Goals:

  • Adopt formal governance structure
  • Recruit additional members (grow to 20-30 organizations)

Activities:

  1. Adopt governance bylaws:
  2. Membership criteria (who can join?)
  3. Voting procedures (one org one vote? weighted?)
  4. Conflict resolution process (mediation, arbitration)

  5. Elect governance board:

  6. Representatives from member organizations
  7. Committees (technical, content policy, ethics, fundraising)

  8. Launch membership drive:

  9. Recruit new members (target: double membership)
  10. Onboarding process (technical setup, governance training)

Phase 4: Scale and Diversify (Years 3-5)

Goals:

  • Grow to 50-100 members
  • Preserve significant collections (multiple murdered platforms)
  • Achieve financial sustainability

Activities:

  1. Expand preservation scope:
  2. Move beyond pilot to major collections
  3. Coordinate triage (use Custodial Filter to prioritize)

  4. Diversify funding:

  5. Membership fees (sliding scale)
  6. Grants (ongoing)
  7. Earned revenue (services to non-members? consulting?)

  8. Build nested structure:

  9. Regional consortia form (US, Europe, Asia, etc.)
  10. Global coordination body emerges
  11. Local autonomy with global standards

Phase 5: Long-Term Sustainability (Years 5+)

Goals:

  • Self-sustaining commons (no dependence on single funder)
  • Recognized as essential infrastructure
  • Continual innovation (technical, governance)

Activities:

  1. Institutionalize:
  2. Become recognized by governments, universities, funders
  3. Partnerships with major institutions (Library of Congress, national libraries)

  4. Adapt and evolve:

  5. Governance reviews (Are Ostrom principles still working?)
  6. Technical upgrades (storage tech, protocols evolve)
  7. Policy updates (new ethical challenges, content types)

  8. Build next generation:

  9. Train new members (how to run nodes, participate in governance)
  10. Document knowledge (guides, case studies, lessons learned)
  11. Inspire new Seed Banks (your model becomes template for others)

Part VI: Why This Is Hard (And How to Succeed Anyway) §

The Challenges §

1. Collective Action Problem

  • Everyone benefits from preservation commons, but contributing is costly
  • Temptation to free-ride (let others do the work)
  • Ostrom’s principles mitigate but don’t eliminate this

2. Technical Complexity

  • Distributed systems are hard to build and maintain
  • Not all organizations have technical capacity
  • Asymmetry in capabilities (big universities vs. small libraries)

3. Funding Uncertainty

  • Commons require sustained funding
  • Grants end, memberships fluctuate
  • Economic downturns threaten budgets

4. Governance Fatigue

  • Democratic governance is labor-intensive (meetings, votes, deliberation)
  • Volunteers burn out
  • Risk of oligarchy (few active members make all decisions)

5. Value Alignment

  • Members must share commitment to commons
  • If some see it as extractive opportunity (monetize data), trust collapses
  • Cultural fit matters—can’t force cooperation

Success Factors §

1. Start Small

  • Don’t try to preserve entire internet on day one
  • Pilot with committed founding members
  • Prove model works, then scale

2. Design Governance First

  • Don’t build tech and add governance later (Mastodon’s mistake)
  • Ostrom’s principles from founding charter
  • Governance evolves but core principles remain

3. Invest in Social Infrastructure

  • Commons succeed when members know and trust each other
  • Regular meetings, workshops, social events
  • Build relationships, not just technical systems

4. Make Contribution Easy

  • Lower barriers to participation (technical documentation, training, support)
  • Graduated membership (start as observer, become full member)
  • Recognize non-technical contributions (curation, governance, outreach)

5. Celebrate Successes

  • Acknowledge contributions publicly
  • Show impact (we preserved X platforms, saved Y terabytes)
  • Build collective pride in commons

6. Be Patient

  • Commons take years to stabilize
  • Early challenges don’t mean failure
  • Ostrom’s principles work but need time

Conclusion: The Seed Bank as Hope §

The Seed Bank isn’t just a technical solution—it’s a political vision. It says:

  • Digital culture doesn’t have to depend on monopolies (Internet Archive is wonderful but shouldn’t be sole preserver)
  • Commons can work (Ostrom proved it, LOCKSS demonstrates it)
  • Cooperation is possible (even in competitive, scarce-resource environments)
  • We can govern ourselves (don’t need corporations or governments to do it for us)

Building a Seed Bank is hard. It requires:

  • Technical expertise (distributed systems, storage, networking)
  • Governance sophistication (Ostrom’s principles aren’t intuitive)
  • Sustained commitment (decades, not months)
  • Cultural alignment (participants must value commons)

But it’s possible. LOCKSS has done it for 20+ years. Other commons can too.

The alternative—centralized preservation monopolies or no preservation at all—is unacceptable. Digital culture is too important, too fragile, too valuable to leave to one organization or to chance.

The Seed Bank is how we preserve digital sovereignty at scale. Not one Archive owned by one organization, but a distributed network of Archives cooperating as a commons.

In the next chapter, we’ll explore the Haunted Forest—how to build memory institutions that don’t just store Umbrabytes, but interpret them, give them meaning, and make them accessible to future generations.

For now, consider: What would it take to start a Seed Bank in your community? Who would you invite? What would you preserve? And how would you govern it together?

The commons begins with an invitation. Will you extend it?


Discussion Questions §

  1. Tragedy Avoided? Ostrom proved Hardin wrong for physical commons (forests, fisheries). Does her work apply to digital commons, or are there fundamental differences?

  2. Trust and Scale: LOCKSS works with 300 libraries. Could it scale to 3,000? 30,000? At what point does commons governance break down?

  3. Mastodon’s Dilemma: Should Mastodon retroactively add governance? Or is pure federation the point (even if it causes problems)?

  4. Your Contribution: If you joined a Seed Bank, what could you contribute? (Technical, financial, labor, curation?) What would you need to participate?

  5. Commons vs. Cooperation: Is a commons (shared resource) better than cooperation between independent archives? What do we gain/lose with commons model?

  6. Nested Governance: How many layers of nesting are optimal? (Node → regional → global? More layers? Fewer?)


Exercise: Design a Seed Bank §

Task: Design a Seed Bank for preserving murdered social media platforms.

Part 1: Apply Ostrom’s Principles (1500 words)

For each of the 8 principles, specify:

  1. How you’ll implement it in your Seed Bank
  2. Specific mechanisms (technical, governance, cultural)
  3. Potential challenges and how to address them

Part 2: Technical Architecture (1000 words)

Choose and justify:

  • Storage model (P2P, federated, hybrid?)
  • Monitoring systems (automated, peer-review, reputation?)
  • Access model (public, restricted, tiered?)
  • Technologies (IPFS, LOCKSS, custom, blockchain, other?)

Part 3: Founding Coalition (500 words)

Who would you recruit as founding members?

  • What types of organizations (universities, libraries, non-profits, individuals?)
  • What commitments would you ask for (storage, funding, labor?)
  • How would you build trust among members?

Part 4: Sustainability Plan (500 words)

How do you keep this going for 20+ years?

  • Funding sources (grants, fees, earned revenue?)
  • Governance evolution (how to prevent ossification or oligarchy?)
  • Technical maintenance (who updates software, migrates formats?)

Part 5: Reflection (300 words)

  • What’s the biggest challenge you anticipate?
  • Would you actually want to participate in this Seed Bank? Why/why not?
  • Is commons model realistic, or too idealistic?

Further Reading §

Elinor Ostrom §

  • Ostrom, Elinor. Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press, 1990.
  • The foundational text—must read

  • Ostrom, Elinor. “Beyond Markets and States: Polycentric Governance of Complex Economic Systems.” American Economic Review 100, no. 3 (2010): 641-672.

  • Nobel Prize lecture, accessible overview

Digital Commons §

  • Benkler, Yochai. The Wealth of Networks: How Social Production Transforms Markets and Freedom. Yale University Press, 2006.
  • Theory of peer production and digital commons

  • Bollier, David, and Silke Helfrich, eds. The Wealth of the Commons: A World Beyond Market and State. Levellers Press, 2012.

  • Case studies of various commons (physical and digital)

  • Hess, Charlotte, and Elinor Ostrom, eds. Understanding Knowledge as a Commons. MIT Press, 2006.

  • Applying commons theory to information/knowledge

LOCKSS and Distributed Preservation §

  • Reich, Vicky, and David S. H. Rosenthal. “LOCKSS: A Permanent Web Publishing and Access System.” D-Lib Magazine 7, no. 6 (2001).
  • Technical overview of LOCKSS system

  • Rosenthal, David S. H. “Emulation & Virtualization as Preservation Strategies.” Report, Library of Congress, 2015.

  • Preservation strategies for distributed systems

Federated Systems and Governance §

  • Gehl, Robert, and Diana Zulli. “Mastodon: Privacy, Moderation, and Affordances of the Private and the Public.” Social Media + Society 5, no. 2 (2019).
  • Analysis of Mastodon’s federated model and challenges

  • Zuboff, Shoshana. The Age of Surveillance Capitalism. PublicAffairs, 2019.

  • Why we need commons alternatives to corporate platforms

End of Chapter 13

Next: Chapter 14 — Memory Institutions for the Digital Age: Curating the Haunted Forest

Part III • Institutional Design, Commons & Political Economy

Chapter 14: Memory Institutions for the Digital Age

Curating the Haunted Forest

24 min read 5,264 words

Opening: The Museum Without Walls §

In 2015, the Strong National Museum of Play in Rochester, New York, opened an exhibit called “World Video Game Hall of Fame.” But this wasn’t a traditional museum exhibit—dusty consoles behind glass with “Do Not Touch” signs. Visitors could play the inducted games: Pong, Pac-Man, Tetris, Doom.

The museum faced a curatorial question that would be absurd in traditional museums: Should we let people touch the artifacts? For video games, the answer had to be yes. A game you can’t play is like a book you can’t read—the medium requires interaction. But interaction means degradation (controllers wear out, CDs get scratched). Museums typically preserve to prevent use. Here, use was preservation—keeping the experience alive.

This is the paradox of digital memory institutions: the artifacts are not just objects to be stored, but experiences to be resurrected. A GeoCities homepage isn’t just HTML files—it’s the experience of navigating a web ring, seeing tags, hearing MIDI music autoplay. Preserving the bits without preserving the context and experience is incomplete.

The Haunted Forest is our metaphor for digital memory institutions—places where murdered platforms and their artifacts exist in liminal space between dead and alive. Not quite functional (the original platform is gone), but not quite inert (the artifacts still haunt us with meaning). Memory institutions curate this haunting—they don’t just store, they interpret, contextualize, and make accessible.

This chapter explores how to build memory institutions for digital culture—museums, archives, libraries, memorials, and research collections that preserve not just bits, but meaning.


Part I: The Five Types of Memory Institutions §

Traditional memory institutions (museums, archives, libraries) each have distinct missions. Digital memory institutions inherit these missions but must adapt them:

Type 1: The Library (Access and Circulation) §

Traditional Mission:

  • Collect materials (books, journals, media)
  • Catalog and organize
  • Lend for temporary use
  • Provide public access

Digital Adaptation: The Web Library

Example: Internet Archive’s Wayback Machine

What it does:

  • Crawls and stores snapshots of websites over time
  • Makes them publicly browsable (800+ billion pages)
  • Free access, no login required
  • Preserves web as if it were a lending library (“borrow” access to past versions)

How it embodies Library mission:

  • Comprehensive collection: Aims to archive “everything” (like Library of Congress)
  • Public access: Anyone can browse, no restrictions (unlike research archives)
  • Findability: URL-based access (like call numbers)
  • Circulation: Multiple users can “use” the same archived page simultaneously

Challenges:

  • Scale is overwhelming (800B pages—impossible to curate comprehensively)
  • Context is minimal (sites preserved but not explained)
  • Robots.txt compliance (respects site owners’ wishes not to be archived—some historically important sites excluded)

When to use Library model:

  • Comprehensive preservation is goal
  • Public access is priority
  • Resources allow for massive scale

Type 2: The Archive (Preservation and Restriction) §

Traditional Mission:

  • Preserve unique/rare materials
  • Maintain original order and provenance
  • Restrict access to protect fragile items
  • Serve researchers, not general public

Digital Adaptation: The Restricted Research Archive

Example: Library of Congress Twitter Archive (2006-2017)

What it does:

  • Preserved all public tweets (billions) from 2006-2017
  • Metadata-only access (can search, but can’t read full tweet text without special permission)
  • Researcher access requires application and justification
  • Not publicly browsable

How it embodies Archive mission:

  • Provenance: Preserves complete record (all tweets, in order, with timestamps)
  • Restriction: Protects privacy (can’t mass-surveil via archive)
  • Research focus: Designed for scholars, not casual browsing
  • Permanence: Committed to preserving forever (unlike platforms)

Challenges:

  • Restrictions limit utility (researchers frustrated by access barriers)
  • Metadata-only access means context is hard to reconstruct
  • 2017 cutoff (stopped collecting—now only selective acquisition)

When to use Archive model:

  • Privacy concerns require restricted access
  • Materials are sensitive or contested
  • Focus is on research, not public engagement

Type 3: The Museum (Display and Interpretation) §

Traditional Mission:

  • Collect objects of cultural/historical significance
  • Curate exhibitions (select and interpret)
  • Educate public through display
  • Create narrative and meaning

Digital Adaptation: The Curated Digital Museum

Example: Cameron’s World (GeoCities Archive as Art)

What it does:

  • Selects GIFs, backgrounds, and UI elements from archived GeoCities sites
  • Arranges them into a sprawling, interactive web collage
  • Provides essays explaining GeoCities aesthetics and culture
  • Makes 1990s web design comprehensible and beautiful

How it embodies Museum mission:

  • Curation: Selects from vast archive (not comprehensive, but meaningful)
  • Interpretation: Explains why GeoCities mattered aesthetically and culturally
  • Exhibition: Public display, visually engaging
  • Education: Teaches people who never experienced GeoCities what it felt like

Another Example: The Strong Museum’s Video Game Hall of Fame

What it does:

  • Inducts significant games into “Hall of Fame” (selective canon)
  • Makes games playable on museum floor (interactive exhibits)
  • Provides historical context (when released, why important, cultural impact)
  • Preserves hardware and software together (full experience)

Challenges:

  • Curation is subjective (who decides what’s significant?)
  • Resources limit scope (can’t exhibit everything)
  • Interpretation can impose narrative (risk of revisionism)

When to use Museum model:

  • Scale is manageable (curated collections, not comprehensive dumps)
  • Public engagement is goal (exhibitions, education)
  • Narrative and interpretation are central

Type 4: The Memorial (Commemoration and Mourning) §

Traditional Mission:

  • Honor the dead or lost
  • Create space for grief and remembrance
  • Preserve memory of trauma or tragedy
  • Offer emotional, not just intellectual, engagement

Digital Adaptation: The Platform Memorial

Example: The September 11 Digital Archive

What it does:

  • Collected personal stories, emails, photos, websites created in response to 9/11
  • Community submissions (people contributed their own materials)
  • Public access (browsable, searchable)
  • Emotional framing (archive as act of collective mourning)

How it embodies Memorial mission:

  • Commemoration: Preserves tragedy and response
  • Personal stories: Not just official record, but individual experiences
  • Emotional resonance: Designed to evoke feeling, not just document facts
  • Community ownership: People participated in creating the archive

Another Example: Hypothetical “GeoCities Memorial”

What it could do:

  • Frame GeoCities shutdown as cultural loss (murder, not natural death)
  • Invite former GeoCities users to submit memories (“Where were you when GeoCities died?”)
  • Create virtual memorial wall (names of lost sites, like Vietnam Memorial)
  • Offer space to grieve the loss of early web’s optimism

Challenges:

  • Emotional framing can seem melodramatic (is a platform shutdown worth mourning?)
  • Risk of nostalgia (romanticizing past at expense of present)
  • Who is being memorialized? (platform? users? era?)

When to use Memorial model:

  • Cultural loss is central (not just preservation, but acknowledging grief)
  • Community needs space to mourn
  • Emotional engagement is goal

Type 5: The Research Collection (Data and Analysis) §

Traditional Mission:

  • Provide raw materials for scholars
  • Emphasis on completeness and accuracy
  • Minimal interpretation (let researchers draw conclusions)
  • Standardized formats for analysis

Digital Adaptation: The Research Dataset

Example: Pushshift Reddit Archive

What it does:

  • Archived every Reddit post and comment (billions) in machine-readable format
  • Made available to researchers (JSON files, searchable API)
  • Minimal curation (raw data dumps)
  • Used for: sociology research, hate speech studies, meme diffusion analysis

How it embodies Research Collection mission:

  • Completeness: Everything archived, not curated sample
  • Machine-readable: JSON, CSV, SQL—formats for computational analysis
  • Researcher-focused: Not public-friendly (requires technical skill)
  • Neutral: Doesn’t interpret data, just provides it

Challenges:

  • Reddit demanded takedown (2023)—Pushshift stopped providing public access
  • Ethical issues: Contains hate speech, harassment, doxxing (should researchers have access?)
  • No interpretation: Requires expertise to use (not accessible to public)

When to use Research Collection model:

  • Scale is massive (too large for manual curation)
  • Goal is to enable research (not public exhibition)
  • Materials are best understood through computational analysis

Part II: The Memory Institution Design Matrix §

When designing a memory institution for murdered digital artifacts, choose your model based on:

Dimension 1: Scale §

Comprehensive (Library/Research Collection)

  • Archive everything or nearly everything
  • Minimal selectivity
  • Example: Internet Archive’s Wayback Machine

Curated (Museum/Memorial)

  • Select significant subset
  • Intensive interpretation
  • Example: Strong Museum’s Video Game Hall of Fame

Dimension 2: Access §

Open (Library/Museum)

  • Public can browse freely
  • No restrictions (or minimal)
  • Example: Cameron’s World, Internet Archive

Restricted (Archive/Research Collection)

  • Requires application, credentials, or justification
  • Protects privacy or sensitivity
  • Example: LOC Twitter Archive

Dimension 3: Interpretation §

High Interpretation (Museum/Memorial)

  • Curators provide context, narrative, meaning
  • Exhibitions tell stories
  • Example: 9/11 Digital Archive with framing essays

Low Interpretation (Archive/Research Collection)

  • Minimal curation, let materials speak for themselves
  • Provenance and metadata, but not narrative
  • Example: Pushshift raw data dumps

Dimension 4: User Experience §

Experiential (Museum)

  • Artifacts are interactive (playable games, browsable sites)
  • Focus on recreating original experience
  • Example: Strong Museum playable games

Documentary (Archive/Library)

  • Artifacts viewed as historical record
  • Screenshots, descriptions, metadata
  • Original experience not replicable
  • Example: Static screenshots of Flash games (not playable)

The Design Matrix §

Institution Type Scale Access Interpretation Experience
Library Comprehensive Open Low Documentary
Archive Comprehensive Restricted Low Documentary
Museum Curated Open High Experiential
Memorial Curated Open High Emotional
Research Collection Comprehensive Restricted Minimal Data-focused

Hybrid Models Are Common:

  • Internet Archive = Library + Archive (comprehensive + some restrictions)
  • Strong Museum = Museum + Research Collection (curated exhibits + comprehensive archives in back)
  • 9/11 Archive = Memorial + Library (emotional framing + open access)

Part III: Curatorial Philosophy — What to Display? §

Museums don’t display everything they own. The Smithsonian’s collections are 95% in storage—only 5% on exhibit. Digital memory institutions face the same question: What do we make visible?

Curatorial Approach 1: Comprehensive Warehouse §

Philosophy: Archive everything, make it all accessible, let users find what they want.

Example: Internet Archive’s Wayback Machine

Strengths:

  • No gatekeeping (curators don’t impose their taste)
  • Serendipity (users discover unexpected things)
  • Completeness (future researchers have maximum material)

Weaknesses:

  • Overwhelming (800B pages—where do you start?)
  • No hierarchy (spam and Shakespeare equally visible)
  • Context is absent (sites preserved without explanation)

Best for: Platforms with structured URLs (websites) where users know what they’re looking for

Curatorial Approach 2: Canon Formation §

Philosophy: Select the “most important” artifacts, create a canon.

Example: Strong Museum’s Video Game Hall of Fame (inducts ~10 games/year)

Strengths:

  • Manageable (visitors can engage deeply with 50 games, not 50,000)
  • Narrative coherence (tells story of video game history)
  • Educational (curators explain why these games matter)

Weaknesses:

  • Elitism (who decides what’s “important”?)
  • Exclusion (marginalizes non-canonical work)
  • Revisionism (canon reflects curator bias)

Best for: Platforms where a small subset represents the whole (pioneering games, influential creators)

Curatorial Approach 3: Thematic Collections §

Philosophy: Organize by themes, movements, or communities.

Example: Hypothetical “Tumblr Fanfiction Archive” organized by fandom, pairing, rating, era

Strengths:

  • Findability (users navigate by interest, not chronology)
  • Contextual (themes provide interpretive frame)
  • Inclusive (multiple themes accommodate diverse interests)

Weaknesses:

  • Subjective (who decides themes?)
  • Overlapping (artifacts fit multiple themes—where do they go?)
  • Incomplete (not everything fits a theme)

Best for: Platforms with identifiable communities or genres (fanfiction, meme culture, activist organizing)

Curatorial Approach 4: Chronological Archive §

Philosophy: Preserve everything in temporal order, like a timeline.

Example: Internet Archive’s snapshots (sites preserved as they changed over time)

Strengths:

  • Objectivity (chronology is neutral)
  • Change visible (see how platforms evolved)
  • Completeness (nothing excluded for thematic reasons)

Weaknesses:

  • No hierarchy (early posts equal to late posts)
  • Narrative absent (time alone doesn’t explain meaning)
  • Scale issues (decades of daily posts = overwhelming)

Best for: Platforms where temporal evolution is key (Twitter’s changing culture, YouTube’s algorithm shifts)

Curatorial Approach 5: Community-Driven Curation §

Philosophy: Let users/creators curate their own materials.

Example: 9/11 Digital Archive (community submissions), Fanlore (fan-created wiki)

Strengths:

  • Authenticity (communities define their own history)
  • Consent (creators choose what’s shared)
  • Diversity (avoids institutional bias)

Weaknesses:

  • Unevenness (some creators participate, others don’t)
  • Coordination challenges (requires infrastructure for submissions)
  • Quality varies (no editorial oversight)

Best for: Platforms where community identity is strong (fandoms, activist movements, hobbyist communities)

Curatorial Approach 6: Algorithmic/Computational Curation §

Philosophy: Use algorithms to select representative samples or identify significant patterns.

Example: Using view counts, shares, replies to identify “most influential” tweets

Strengths:

  • Scalability (algorithms process massive datasets)
  • Objectivity (no human bias—though algorithms have bias too)
  • Discovery (finds patterns humans miss)

Weaknesses:

  • Black box (users don’t know why things were selected)
  • Bias (algorithms reflect creator bias and training data)
  • Misses margins (algorithms favor mainstream)

Best for: Platforms with clear metrics (views, likes, shares) and massive scale


Part IV: Technical Fidelity — How Much to Preserve? §

Digital artifacts exist in layers. How much of each layer do you preserve?

The Fidelity Ladder §

Level 1: Documentation Only

What’s preserved: Screenshots, descriptions, metadata

What’s lost: Interactivity, experience, technical details

Example: Wikipedia article about Vine (describes it, but can’t show it)

Pros: Cheap, easy, lightweight Cons: Least faithful to original

When to use: Platform is already dead, no way to preserve fully; documentation better than nothing

Level 2: Static Archive

What’s preserved: HTML, CSS, images (rendered as static files)

What’s lost: JavaScript interactivity, dynamic content, databases

Example: Archived GeoCities sites (HTML works, but embedded widgets/scripts don’t)

Pros: Relatively easy, preserves visual appearance Cons: Non-interactive sites feel “dead”

When to use: Static websites, blogs, simple HTML pages

Level 3: Emulation

What’s preserved: Full functionality via emulator (browser, OS, hardware)

What’s lost: Original hardware experience (speed, bugs, quirks)

Example: Flash games playable via Ruffle emulator, DOS games via DOSBox

Pros: Fully interactive, close to original experience Cons: Requires maintaining emulators (which can become obsolete)

When to use: Complex platforms requiring specific environments (Flash, Java, old browsers)

Level 4: Source Code Preservation

What’s preserved: Actual code, databases, server configurations

What’s lost: Nothing (in theory)—but requires technical expertise to run

Example: GitHub archives of open-source projects

Pros: Most faithful, can be recompiled/forked/modified Cons: Requires developer skills, dependencies may be obsolete

When to use: Open-source platforms, when preserving for future developers (not just users)

Level 5: Live Preservation

What’s preserved: Original infrastructure still running

What’s lost: Nothing (it’s still alive)

Example: Old arcade games kept running on original hardware by collectors

Pros: Perfect fidelity Cons: Expensive, fragile (hardware fails), not scalable

When to use: High-value artifacts where experience depends on specific hardware (rare)

Level 6: Resurrection

What’s preserved: Platform rebuilt from scratch for modern environments

What’s lost: Bugs, quirks, historical authenticity (new code ≠ old code)

Example: Homestar Runner rebuilt in HTML5 (originally Flash)

Pros: Accessible on modern devices, no emulation needed Cons: Not “authentic” (it’s a recreation, not preservation)

When to use: Cultural value is high, original platform can’t run anymore, resurrection is only option

The Fidelity Trade-off §

Higher fidelity = higher cost (time, storage, maintenance, expertise)

Strategy: Tiered preservation

  • Level 1-2 (documentation/static): Archive everything
  • Level 3-4 (emulation/source): Archive high-value subset
  • Level 5-6 (live/resurrection): Only most culturally significant artifacts

Part V: Access and Discovery — Making the Haunted Forest Navigable §

Preserving artifacts is half the battle. Making them findable and usable is the other half.

Access Model 1: URL-Based (Library Model) §

How it works: Every artifact has a permanent URL; users navigate directly or via search engines

Example: Internet Archive’s Wayback Machine (web.archive.org/web/TIMESTAMP/URL)

Pros:

  • Simple, intuitive
  • Integrates with web (can link from anywhere)
  • Decentralized (no need for central index)

Cons:

  • Requires knowing URL (hard if you don’t remember the site)
  • No thematic browsing (can’t explore by topic)

Access Model 2: Search-Based (Database Model) §

How it works: Full-text search across all preserved content

Example: Archive.org’s search bar, Google Books

Pros:

  • Powerful discovery (find anything containing keyword)
  • Don’t need to know exact URL

Cons:

  • Overwhelming (thousands of results)
  • Poor for browsing (good for finding specific thing, bad for exploration)

Access Model 3: Curated Exhibits (Museum Model) §

How it works: Curators create thematic collections or virtual exhibitions

Example: Strong Museum’s Hall of Fame induction pages, Cameron’s World

Pros:

  • Guided experience (learn through narrative)
  • Manageable scope (100 items, not 100,000)
  • Contextual (exhibits explain significance)

Cons:

  • Limited (most collection not exhibited)
  • Curator bias (what’s not exhibited is invisible)

Access Model 4: Community Wikis (Collaborative Model) §

How it works: Community members add metadata, tags, context

Example: Fanlore (fan-created wiki about fandom history), Wikipedia’s coverage of internet culture

Pros:

  • Distributed labor (community shares work)
  • Insider knowledge (fans know context outsiders miss)
  • Self-updating (as community learns, wiki improves)

Cons:

  • Uneven coverage (popular fandoms well-documented, niche ones ignored)
  • Quality varies (no editorial oversight)
  • Requires active community (if community dies, wiki stagnates)

Access Model 5: API-Based (Researcher Model) §

How it works: Machine-readable access (JSON, CSV, SQL) for computational analysis

Example: Pushshift API, Twitter Academic API

Pros:

  • Enables large-scale research (computational humanities, data science)
  • Flexible (researchers query exactly what they need)

Cons:

  • Not user-friendly (requires programming skill)
  • Not browsable (can’t casually explore)

Hybrid Access Strategy §

Most memory institutions use multiple access methods:

  • URLs for direct access (if you know what you want)
  • Search for discovery (find specific content)
  • Exhibits for education (learn about the platform/era)
  • Wiki for context (community-generated metadata)
  • API for research (scholars analyze at scale)

Example: Internet Archive

  • Wayback Machine: URL-based
  • Search bar: keyword search
  • Collections: curated thematic groups (e.g., “Grateful Dead Live Concerts”)
  • API: developers can query programmatically

Memory institutions must navigate thorny legal and ethical issues:

Problem: Most preserved content is copyrighted. Does archiving violate copyright law?

Legal Frameworks:

Fair Use (US)

  • Preservation may qualify as fair use (transformative, educational, minimal market harm)
  • Case law: Authors Guild v. Google (Google Books scanning ruled fair use)
  • But: Fair use is defense, not right—you could still be sued

Section 108 (US Copyright Law)

  • Libraries and archives can preserve copyrighted works under certain conditions:
  • Non-commercial purpose
  • Closed systems (access in library only, or limited digital access)
  • But: Doesn’t cover web scraping or mass digitization clearly

DMCA Safe Harbor

  • Platforms not liable for user-uploaded content if they respond to takedowns
  • Memory institutions use this (Internet Archive responds to DMCA requests)

International Variations:

  • EU: Orphan Works Directive allows preservation of works with unknown copyright holders
  • Canada: Fair Dealing (narrower than US fair use)

Practical Strategy:

  • Preserve first, respond to takedowns if challenged (Internet Archive’s approach)
  • Or: Seek permissions (time-consuming, often impossible)
  • Or: Restrict access (preserve but don’t make public)

Issue 2: Privacy §

Problem: Archived content may contain personal information people no longer want public.

Ethical Questions:

  • Should you preserve someone’s teenage LiveJournal (they might be embarrassed now)?
  • Should you archive doxxing or harassment (evidence of harm, but re-publicizes victim info)?
  • Should you preserve medical/financial/intimate details shared on forums?

Frameworks:

Right to Be Forgotten (GDPR, EU)

  • Individuals can request deletion of personal data
  • Applies to archives? Unclear (exemptions for journalism/research/public interest)

Contextual Integrity (Helen Nissenbaum)

  • Privacy violated when information flows across contexts inappropriately
  • Example: Forum post meant for small community, now archived and Google-indexed = context collapse

Practical Approaches:

Takedown Policies:

  • Allow individuals to request removal (Internet Archive honors requests)
  • Review case-by-case (balance individual privacy vs. historical value)

Restricted Access:

  • Preserve but don’t make publicly searchable
  • Researcher-only access (requires IRB approval)

Anonymization:

  • Remove or redact names, usernames, identifying details
  • But: Can harm historical accuracy

Problem: Did creators consent to their work being preserved?

Arguments:

Implied Consent:

  • By posting publicly, you consented to archiving (like publishing a book)
  • Counterargument: Expectation of ephemerality (platform may shut down, but users didn’t expect Internet Archive)

Explicit Consent:

  • Only archive if creator explicitly agrees
  • Counterargument: Impractical (can’t contact millions of users)

Posthumous:

  • If creator is dead, do we need consent from estate?
  • Historical materials often preserved without consent (diaries, letters found after death)

Practical Strategy:

  • Default to preserving (implied consent for public posts)
  • Honor explicit deletion (if someone deleted content, don’t resurrect without reason)
  • Provide opt-out (let creators request removal)

Issue 4: Harm Prevention §

Problem: Some content causes harm if preserved (hate speech, doxxing, revenge porn).

Ethical Framework:

Do No Harm Principle:

  • If preserving causes direct, ongoing harm (reveals someone’s address, enables harassment), don’t do it

Historical Value vs. Harm:

  • Hate forums: preserve for research (understanding extremism), but restrict access (don’t make recruitment tool)
  • Revenge porn: don’t preserve (no historical value justifies harm)

Contextual Judgment:

  • Evaluate case-by-case
  • Consult affected communities when possible

Part VII: Case Studies in Memory Institution Design §

Case Study 1: The Strong Museum (Exemplary Museum Model) §

What they do:

  • Curate exhibitions of video games, toys, and play
  • Make games playable (interactive exhibits)
  • Preserve hardware and software together
  • Host researchers (extensive archives beyond exhibits)

Why it works:

  • Curation: Selective canon (Hall of Fame inductees)
  • Experience: Games are played, not just viewed
  • Interpretation: Context provided (essays, talks, labels)
  • Institutional stability: Endowed museum (not dependent on platform survival)

Challenges:

  • Limited scope (only games, not all digital culture)
  • Geography-bound (must visit Rochester to play games)

Case Study 2: Internet Archive (Exemplary Library Model) §

What they do:

  • Crawl and preserve websites (Wayback Machine)
  • Archive books, music, video, software
  • Open access (free, no login)
  • Advocate for digital rights (lawsuits for fair use)

Why it works:

  • Scale: 800+ billion web pages
  • Longevity: 29 years and counting
  • Public good: Non-profit, funded by donations and services
  • Legal courage: Willing to defend fair use in court

Challenges:

  • Scale makes curation impossible (overwhelming)
  • Robots.txt compliance excludes important sites
  • Funding precarity (dependent on donations)

Case Study 3: Fanlore (Exemplary Community-Driven Model) §

What they do:

  • Wiki documenting fandom history (ships, tropes, controversies, communities)
  • Created and maintained by fans
  • Covers all fandoms (TV, books, games, RPF, etc.)

Why it works:

  • Insider knowledge: Fans document nuances outsiders miss
  • Community ownership: Fans preserve their own history
  • Distributed labor: Thousands of contributors

Challenges:

  • Uneven coverage (big fandoms well-documented, small ones sparse)
  • Vandalism/edit wars (controversial topics fought over)
  • Succession (if volunteer community dwindles, wiki could die)

Case Study 4: Flashpoint Project (Exemplary Resurrection Model) §

What they do:

  • Preserve 500,000+ Flash games and animations
  • Provide custom launcher with embedded emulator
  • Community-curated (volunteers add games, metadata)

Why it works:

  • Rescue mission: Saved massive amount of content before Flash died
  • Playability: Games fully functional (not just archived)
  • Community-driven: Distributed effort (volunteers curate, test, tag)

Challenges:

  • Maintenance burden (emulators need updates as OSes change)
  • Copyright gray area (hosting games without explicit permission)
  • Curation slow (500K games, but millions more exist—can’t save everything)

Part VIII: Building Your Own Memory Institution §

Step-by-Step Guide §

Phase 1: Define Mission

Questions:

  1. What are you preserving? (specific platform, genre, community, era)
  2. Why does it matter? (cultural significance, underrepresentation, endangerment)
  3. Who is your audience? (general public, researchers, community members)
  4. What type of institution? (library, archive, museum, memorial, research collection)

Example Mission: “The Tumblr Fandom Archive preserves fanworks (fanfiction, fan art, meta) from Tumblr’s golden age (2010-2016), focusing on marginalized fandoms and LGBTQ+ creators. Our audience is fans, scholars, and future generations interested in transformative works. We are a community-driven digital museum with curated exhibits and open archives.”

Phase 2: Acquisition Strategy

How will you acquire content?

Option A: Scrape

  • Use tools (wget, ArchiveBox, Heritrix)
  • Pros: Comprehensive
  • Cons: Legal gray area, may violate ToS

Option B: Community Submissions

  • Invite creators to submit their work
  • Pros: Consent-based, community-driven
  • Cons: Incomplete (only those who participate)

Option C: Partnerships

  • Work with platform for data dump
  • Pros: Legal, comprehensive
  • Cons: Requires cooperation (rare)

Option D: Hybrid

  • Scrape public content + invite submissions for private/deleted content

Phase 3: Storage and Infrastructure

Technical Needs:

  • Storage: Servers or cloud (how much data?)
  • Redundancy: Multiple backups (LOCKSS principle)
  • Access platform: Website for browsing (static site, database-driven, CMS)

Options:

  • Self-hosted: Full control, but maintenance burden
  • Cloud: Scalable, but ongoing costs
  • Partner with institution: University library hosts (stability, but less autonomy)

Budget:

  • Small project: $100-500/year (domain, shared hosting, cloud storage)
  • Medium project: $5K-50K/year (dedicated servers, staff time)
  • Large project: $100K+/year (institutional scale, like Internet Archive)

Phase 4: Curation and Metadata

How will you organize content?

Metadata Schema:

  • Dublin Core (standard for libraries)
  • Custom schema (specific to your domain)

Essential Fields:

  • Title, Creator, Date, Description, Tags/Categories, Source URL, Archive Date

Curation Approach:

  • Comprehensive warehouse (minimal curation)
  • Thematic collections (curated exhibits)
  • Algorithmic (automated tagging, recommendations)
  • Community-driven (user-submitted metadata)

Phase 5: Access and Discovery

How will users find content?

Build:

  • Search functionality (full-text or metadata)
  • Browse by category, date, creator
  • Featured/curated collections (homepage highlights)
  • API (for researchers)

Tools:

  • Static site generator (Jekyll, Hugo) for simple projects
  • Database + CMS (WordPress, Omeka) for complex projects
  • Custom web app (Flask, Django, Rails) for maximum control

Document:

  • Copyright stance (fair use, takedown policy)
  • Privacy policy (what personal info do you collect/preserve?)
  • Consent framework (do you allow opt-out?)
  • Access restrictions (public, researcher-only, embargoed)

Get advice:

  • Consult copyright lawyer
  • Follow models (Internet Archive’s policies, university IRB guidelines)

Phase 7: Launch and Maintenance

Launch:

  • Soft launch (invite community, gather feedback)
  • Public announcement (blog post, social media, press)

Ongoing:

  • Add content regularly (don’t let it stagnate)
  • Respond to takedown requests
  • Update software/emulators (prevent bit rot)
  • Fundraise (donations, grants, sponsorships)

Succession Planning:

  • What happens if you can’t maintain it? (partner institution, hand off to community, deposit in Internet Archive)

Conclusion: Curating Haunted Spaces §

Memory institutions for digital culture are not just storage facilities—they’re acts of interpretation. Every curatorial choice (what to preserve, how to display, who gets access) shapes how future generations understand our present.

The Haunted Forest is haunted precisely because these artifacts are liminal—neither alive nor fully dead. They exist in the gap between platform death and historical canonization. Memory institutions curate this gap, transforming murdered platforms into ghosts that can teach, inspire, and warn.

When you build a memory institution, you’re not just saving bits. You’re:

  • Resurrecting experience (making dead platforms playable, browsable, meaningful)
  • Creating context (explaining why this mattered, what it meant, who it served)
  • Enabling research (providing materials for scholars)
  • Honoring loss (memorializing what platforms murdered)
  • Building canon (deciding what future remembers)

The question isn’t just “Can we preserve this?” but “How do we make this meaningful for people who never experienced it?”

In the next chapter, we move from memory institutions to political economy—examining the Sovereignty Stack and how to redesign the infrastructure that platforms control.

But first, go build a memory institution. Even a small one. Preserve something meaningful to you. Curate it. Interpret it. Make it accessible.

The Haunted Forest needs its curators.


Discussion Questions §

  1. Institutional Identity: If you were building a memory institution for a murdered platform, which model (library, archive, museum, memorial, research collection) would you choose? Why?

  2. Curation vs. Comprehensiveness: Should memory institutions try to preserve everything, or curate selectively? What are the ethical stakes of each approach?

  3. Fidelity Trade-offs: How much technical fidelity is “enough”? When is a screenshot sufficient vs. needing full emulation?

  4. Access Politics: Who should have access to preserved materials? Public? Researchers only? Community members only? How do you balance openness with privacy?

  5. Canon Formation: Who decides what’s “historically significant”? How do we avoid reproducing bias in digital preservation?

  6. Your Own Archive: What digital artifact from your life would you want preserved in a memory institution? How would you want it curated and displayed?


Exercise: Design a Memory Institution §

Scenario: Choose a platform that has died or is dying (MySpace, Vine, GeoCities, Google+, Tumblr’s NSFW content, etc.). Design a memory institution to preserve and present it.

Part 1: Mission and Model (500 words)

  • What’s your institution called?
  • What type (library, archive, museum, memorial, research collection, hybrid)?
  • What’s your mission statement?
  • Who is your audience?

Part 2: Collection Strategy (500 words)

  • What will you preserve? (everything, curated subset, specific communities)
  • How will you acquire it? (scraping, submissions, partnership)
  • What’s the scale? (how much content, storage needed)

Part 3: Curatorial Approach (500 words)

  • How will you organize content? (comprehensive warehouse, thematic collections, chronological, community-driven)
  • What metadata will you capture?
  • What level of technical fidelity? (documentation, static, emulation, source code, live)

Part 4: Access and Discovery (500 words)

  • How will users find content? (URL-based, search, exhibits, wiki, API)
  • What’s publicly accessible vs. restricted?
  • How will you handle privacy/consent/copyright?

Part 5: Implementation Plan (500 words)

  • What infrastructure do you need? (servers, storage, software)
  • Budget estimate (startup + annual maintenance)
  • Staffing (who does what? volunteers or paid?)
  • Sustainability plan (how do you keep it running for 50 years?)

Part 6: Reflection (300 words)

  • What’s the biggest challenge?
  • What compromises did you make (fidelity vs. budget, comprehensiveness vs. curation, access vs. privacy)?
  • Would you actually want to build this? Why or why not?

Further Reading §

On Museums and Memory §

  • Kirshenblatt-Gimblett, Barbara. Destination Culture: Tourism, Museums, and Heritage. University of California Press, 1998.
  • Young, James. The Texture of Memory: Holocaust Memorials and Meaning. Yale University Press, 1993.
  • Crane, Susan. “Memory, Distortion, and History in the Museum.” History and Theory 36, no. 4 (1997): 44-63.

On Digital Curation §

  • Manoff, Marlene. “Archive and Database as Metaphor: Theorizing the Historical Record.” Portal: Libraries and the Academy 10, no. 4 (2010): 385-398.
  • Brügger, Niels. “Website History and the Website as an Object of Study.” New Media & Society 11, no. 1-2 (2009): 115-132.
  • Owens, Trevor. The Theory and Craft of Digital Preservation. Johns Hopkins University Press, 2018.

On Interpretation and Exhibition §

  • Hooper-Greenhill, Eilean. Museums and the Interpretation of Visual Culture. Routledge, 2000.
  • Macdonald, Sharon, ed. A Companion to Museum Studies. Wiley-Blackwell, 2011.
  • Pearce, Susan. Museums, Objects and Collections. Leicester University Press, 1992.

On Video Game Preservation §

  • Newman, James. Best Before: Videogames, Supersession and Obsolescence. Routledge, 2012.
  • Guttenbrunner, Mark, et al. “Keeping the Game Alive: Evaluating Strategies for the Preservation of Console Video Games.” International Journal of Digital Curation 5, no. 1 (2010): 64-90.

Case Studies §

  • Internet Archive. “About the Internet Archive.” https://archive.org/about/
  • The Strong Museum. “About the Strong.” https://www.museumofplay.org/about/
  • Fanlore. “About Fanlore.” https://fanlore.org/wiki/Fanlore:About
  • Flashpoint Project. https://flashpointarchive.org/

End of Chapter 14

Next: Part IV — Systems & Movements Chapter 15 — The Political Economy of Digital Ground

Part III • Institutional Design, Commons & Political Economy

Chapter 15: The Political Economy of Digital Ground

Who Controls the Infrastructure?

22 min read 4,804 words

Opening: The Day the Domain Disappeared §

In 2010, the United States government seized the domain mooo.com—a URL shortener popular with music fans and file sharers. Without warning, without trial, without due process, the Department of Homeland Security simply took control of the domain and replaced the site with a banner: “This domain name has been seized by ICE—Homeland Security Investigations.”

Thousands of links across the internet broke instantly. Blog posts, forum threads, social media shares—all pointed to dead URLs. The content those links pointed to still existed on other servers, but the naming system had been captured. The government didn’t need to touch the actual files; they just seized the address that pointed to them.

This wasn’t an isolated incident:

  • Libya (.ly domains, 2010): Shut down vb.ly (URL shortener) for “violating Islamic morality”
  • Kazakhstan (.kz, 2021): Temporarily seized opposition media domains during protests
  • Ukraine (.ua, 2014): Domains hijacked during Crimea annexation
  • China (.cn, ongoing): Routine domain seizures for political speech

The lesson: Even if you own your content, if you don’t control the infrastructure that makes it accessible, you don’t truly have Ground.

This chapter explores the political economy of digital infrastructure—who controls the layers that make the internet work, how that control is exercised, and how we might redesign those layers to resist capture.

We’ll introduce the Sovereignty Stack: six layers of digital infrastructure, each with different ownership models and vulnerabilities. Then we’ll examine case studies of each layer being captured or contested. Finally, we’ll explore alternatives being built to resist centralized control.


Part I: The Sovereignty Stack — Six Layers of Digital Infrastructure §

Digital sovereignty isn’t just about owning a domain or hosting a website. It requires control (or at least resilience) across six interconnected layers:

Layer 1: Physical Infrastructure (Bottom Layer) §

What it is:

  • Undersea cables carrying internet traffic between continents
  • Data centers housing servers
  • Cell towers and fiber optic lines
  • Electricity grids powering everything

Who controls it:

  • Telecom corporations (AT&T, Comcast, China Telecom)
  • Cloud providers (Amazon AWS, Microsoft Azure, Google Cloud)
  • Nation-states (can cut cables, seize data centers)

Sovereignty implications:

  • If you don’t own physical hardware, you’re renting from someone who can evict you
  • Governments can surveil traffic at chokepoints (NSA’s undersea cable taps)
  • Interruptions (power outages, cable cuts) can take down entire regions

Vulnerability:

  • Centralization: Most cloud traffic flows through a few companies’ data centers
  • Geographic concentration: Major cable landing points are in a few cities
  • State power: Physical infrastructure is ultimately subject to whoever controls territory

Resistance strategies:

  • Distributed hosting (content on multiple continents)
  • Peer-to-peer networks (no central servers)
  • Community-owned ISPs (fiber cooperatives)
  • Mesh networks (neighborhood-scale wireless networks)

Layer 2: Network Protocols (Transport Layer) §

What it is:

  • TCP/IP (how data packets move across networks)
  • HTTP/HTTPS (how browsers talk to servers)
  • DNS (Domain Name System—translates names like google.com to IP addresses)
  • BGP (Border Gateway Protocol—routes traffic between networks)

Who controls it:

  • Standards bodies (IETF, W3C—mostly open)
  • ICANN (manages DNS root servers)
  • ISPs and network operators (implement protocols)

Sovereignty implications:

  • Open protocols (like TCP/IP) are more sovereign than proprietary ones
  • DNS is a single point of failure—if your domain is seized, your site becomes unreachable
  • BGP hijacking can redirect traffic (happened to YouTube, Amazon, others)

Vulnerability:

  • DNS centralization: ICANN ultimately controls the root DNS servers
  • Certificate authorities: HTTPS requires CAs, which can be compromised or coerced
  • Routing attacks: BGP has no built-in authentication (trust-based)

Resistance strategies:

  • Alternative DNS (blockchain-based names like ENS, Namecoin)
  • Onion routing (Tor—routes traffic through multiple nodes to hide destination)
  • IPFS (content-addressed instead of location-addressed—content hash is the “address”)
  • Mesh routing (packets find paths dynamically, no central routing tables)

Layer 3: Identity Systems (Authentication Layer) §

What it is:

  • How you prove you are who you claim to be
  • Username/password on platforms
  • Email addresses
  • Social login (“Sign in with Google/Facebook”)
  • Digital certificates and keys

Who controls it:

  • Platforms (Facebook, Google, Apple control their login systems)
  • Federated identity providers (OAuth services)
  • Individuals (if using self-hosted identity)

Sovereignty implications:

  • Platform identity is leased (Facebook can delete your account, erasing your identity)
  • Email is more sovereign (you can change providers, keep your address if you own domain)
  • Federated identity (like Mastodon’s @[email protected]) gives you control

Vulnerability:

  • Platform capture: Most people use “Sign in with Google/Facebook” (convenient but gives those companies control)
  • Real name policies: Platforms requiring legal names harm pseudonymous freedom
  • Account suspension: Losing your account means losing your identity across all connected services

Resistance strategies:

  • Self-hosted identity (your own domain, your own email server)
  • Decentralized identifiers (DIDs—cryptographic identities not tied to any platform)
  • PGP/GPG keys (cryptographic proof of identity)
  • Federated identity (ActivityPub, IndieAuth)

Layer 4: Data Storage (Persistence Layer) §

What it is:

  • Where your files, messages, posts, photos actually live
  • Cloud storage (Google Drive, Dropbox, iCloud)
  • Platform databases (Facebook’s servers storing your posts)
  • Self-hosted storage (your own hard drive or server)

Who controls it:

  • Cloud providers (Amazon S3, Google Cloud Storage)
  • Platforms (Twitter, Instagram, TikTok)
  • Individuals (if self-hosting)

Sovereignty implications:

  • If your data lives on someone else’s servers, they can delete it, surveil it, or lock you out
  • Export tools help (you can download your data), but migrating is often hard
  • Self-hosting gives you control but requires technical skill and maintenance

Vulnerability:

  • Terms of Service changes: Provider can change terms, delete your data, or raise prices
  • Platform shutdown: If the company dies, your data dies (unless you exported)
  • Vendor lock-in: Proprietary formats make it hard to migrate

Resistance strategies:

  • Self-hosting (Nextcloud, Syncthing)
  • Distributed storage (IPFS, BitTorrent, Filecoin)
  • Local-first software (data lives on your device, syncs peer-to-peer)
  • Regular exports (even if using cloud, keep local backups)

Layer 5: Application Layer (Interface Layer) §

What it is:

  • The software you actually use: social media apps, email clients, browsers, editors
  • Web apps (run in browser, on company servers)
  • Native apps (run on your device, may or may not depend on servers)
  • Protocols (open standards that any app can implement)

Who controls it:

  • Platform companies (Facebook, Twitter, TikTok)
  • Open source communities (Firefox, Linux, WordPress)
  • Standards bodies (W3C for web standards, IETF for internet protocols)

Sovereignty implications:

  • Proprietary apps lock you into a platform’s ecosystem
  • Open source apps can be forked if the company sells out
  • Open protocols allow multiple competing apps (email: Gmail, Outlook, Apple Mail all work together)

Vulnerability:

  • Platform changes: Twitter can redesign its app, remove features users rely on
  • API shutdowns: Third-party apps can be killed (Twitter banned third-party clients in 2023)
  • Dark patterns: Apps designed to be addictive, manipulative (infinite scroll, notification spam)

Resistance strategies:

  • Use open source apps (can’t be taken away)
  • Use protocol-based tools (email, RSS, ActivityPub)
  • Support third-party clients (don’t let platforms monopolize access)
  • Build alternatives (if platform sucks, fork it or build competitor)

Layer 6: Economic Layer (Value Layer) §

What it is:

  • How money flows through digital systems
  • Payment processors (Visa, PayPal, Stripe)
  • Platform monetization (ads, subscriptions, cuts of transactions)
  • Cryptocurrency (Bitcoin, Ethereum)
  • Creator economy (Patreon, Ko-fi, OnlyFans)

Who controls it:

  • Payment oligopolies (Visa/Mastercard process ~80% of transactions)
  • Platform companies (take 30% cuts on app stores, 15-50% on creator platforms)
  • Banks and regulators (can freeze accounts, block payments)

Sovereignty implications:

  • If platforms or payment processors can cut off your income, you’re not economically sovereign
  • Deplatforming often includes payment bans (see: sex workers, political dissidents, WikiLeaks)
  • Creator economy gives some independence, but platforms still take large cuts

Vulnerability:

  • Payment processor power: Visa/PayPal can ban you, cutting off all income
  • Platform rent-seeking: App stores take 30%, creator platforms take 15-50%
  • Regulatory capture: Governments can pressure payment companies to ban disfavored actors

Resistance strategies:

  • Cryptocurrency (censorship-resistant payments, but volatile and technical barriers)
  • Direct payments (checks, cash, wire transfers—cumbersome but sovereign)
  • Cooperatively-owned payment systems (credit unions, payment co-ops)
  • Multiple revenue streams (don’t depend on one platform or processor)

Part II: The Stack in Practice — Case Studies of Control and Resistance §

Case Study 1: DNS Control — The .ly Seizure (Layer 2) §

What Happened:

In 2010, Libya (which controls the .ly top-level domain) seized vb.ly, a popular URL shortener, for “violating Libyan Islamic morality laws.” The site had shortened links to content Libya deemed offensive.

Impact:

  • Thousands of shortened URLs broke instantly
  • Blogs, tweets, forum posts—all pointing to dead links
  • Content wasn’t deleted, but became unreachable via those URLs

Sovereignty Stack Analysis:

Layer Vulnerability Lesson
Physical Content was fine (on US servers) Not the issue
Network (DNS) CRITICAL FAILURE Libya controlled .ly TLD, seized the domain
Identity Vb.ly users lost identity (their short URLs) Dependent on DNS
Storage Content still existed But unreachable
Application URL shortener app still worked But domain gone
Economic Vb.ly lost revenue Revenue depends on domain access

Lesson: Even if you control Layers 4-6 (storage, apps, payment), if Layer 2 (DNS) is captured, you’re toast.

Resistance: Use domains under TLDs controlled by stable, rights-respecting jurisdictions. Or use alternative naming (onion addresses, IPFS hashes, blockchain names).

Case Study 2: AWS Deplatforming — Parler (Layer 1) §

What Happened:

In January 2021, Amazon Web Services (AWS) terminated Parler’s hosting contract, citing Terms of Service violations (insufficient moderation of violent content). Parler went offline for a month until they found alternative hosting.

Impact:

  • Entire platform inaccessible (no hosting = no site)
  • Millions of users locked out
  • Parler eventually returned via smaller hosting providers, but with degraded performance

Sovereignty Stack Analysis:

Layer Vulnerability Lesson
Physical CRITICAL FAILURE AWS controlled servers, kicked Parler off
Network Parler still had their domain But no servers to point to
Identity User accounts existed (in database) But database offline
Storage Data existed (Parler had backups) But no public access
Application Parler’s app still existed But servers offline
Economic Couldn’t run ads or process payments Business model dead

Lesson: If you don’t own Layer 1 (physical infrastructure), you’re vulnerable to hosting providers’ decisions. AWS, Google Cloud, and Azure are de facto gatekeepers.

Resistance: Self-host (expensive and technical), use offshore hosting (politically risky), or distribute across multiple providers (complex but resilient).

Case Study 3: Platform Identity Control — Facebook’s Real Name Policy (Layer 3) §

What Happened:

Facebook required users to use their legal names (2014 policy enforcement wave). Thousands of accounts suspended:

  • Drag performers using stage names
  • Native Americans with non-Western naming conventions
  • Abuse survivors hiding from stalkers
  • LGBTQ+ people using chosen names

Impact:

  • Users lost access to years of posts, photos, friend networks
  • Some lost business pages tied to stage personas
  • Forced outing for trans people using new names

Sovereignty Stack Analysis:

Layer Vulnerability Lesson
Physical Not the issue Servers worked fine
Network Not the issue DNS worked fine
Identity CRITICAL FAILURE Facebook controlled identity, could revoke it
Storage Data existed but locked Users couldn’t access their own content
Application Facebook app worked But identity removed
Economic Lost business pages Revenue tied to identity

Lesson: Platform-controlled identity is leased, not owned. If the platform can delete your identity, you have no Declaration (Pillar 1).

Resistance: Use federated identity (@[email protected]), self-host identity, use cryptographic keys (not platform usernames).

Case Study 4: Payment Processor Deplatforming — WikiLeaks (Layer 6) §

What Happened:

In 2010, after WikiLeaks published classified US documents, Visa, Mastercard, PayPal, and Bank of America all cut off WikiLeaks’ payment processing. Donations plummeted 95%.

Impact:

  • WikiLeaks nearly went bankrupt
  • Couldn’t accept credit card donations
  • Had to rely on Bitcoin (before it was widely adopted)

Sovereignty Stack Analysis:

Layer Vulnerability Lesson
Physical Hosting providers also pressured But WikiLeaks migrated
Network Domain pressured but survived DNS remained
Identity Not the issue WikiLeaks identity intact
Storage Archives remained accessible Data fine
Application Website worked Accessible
Economic CRITICAL FAILURE Payment processors cut off revenue

Lesson: Economic sovereignty (Layer 6) is essential. If payment processors can deplatform you, you’re economically vulnerable no matter how technically sovereign you are.

Resistance: Bitcoin (WikiLeaks adopted it, later profited from appreciation), direct bank transfers (slow, high fees), cash/checks (physical, not scalable).

Case Study 5: Mastodon Federation — Distributed Resilience (Layers 2-5) §

What Happened:

Mastodon launched in 2016 as a federated alternative to Twitter. Anyone can run a Mastodon instance (server); instances communicate via ActivityPub protocol.

Resilience across layers:

Layer How Mastodon Resists Capture
Physical Distributed servers (no single company owns all hosting)
Network Open protocol (ActivityPub—any server can join)
Identity Federated (@[email protected]—tied to instance, but migratable)
Storage Instance admin controls data (users can export, migrate)
Application Open source (anyone can fork, improve, or run their own version)
Economic Varied models (donations, subscriptions, volunteer-run)

Advantages:

  • No single point of failure (if one instance dies, others survive)
  • No corporate control (community-governed instances)
  • Portable identity (can migrate between instances)

Challenges:

  • Instance admin power (can still ban you from that instance)
  • Fragmentation (instances defederate, creating silos)
  • Technical barriers (running an instance requires skill)
  • Economic sustainability (many instances struggle to fund themselves)

Lesson: Federation distributes sovereignty but doesn’t eliminate all vulnerabilities. Still better than centralized platforms.


Part III: Redesigning the Stack — Alternatives to Centralized Control §

Layer 1 Alternatives: Physical Infrastructure §

Problem: Cloud oligopoly (AWS, Google Cloud, Azure control most hosting)

Alternatives:

Community-Owned ISPs:

  • Fiber cooperatives (users own the network)
  • Examples: Chattanooga’s municipal fiber, NYC Mesh
  • Advantages: Democratic control, no corporate extraction
  • Challenges: Requires local organizing, capital investment

Peer-to-Peer Hosting:

  • IPFS (InterPlanetary File System): Content distributed across many nodes
  • BitTorrent: Files seeded by many users, no central server
  • Advantages: No single point of failure, censorship-resistant
  • Challenges: Slow for dynamic content, legal gray areas

Mesh Networks:

  • Wireless networks where users’ devices relay traffic
  • Example: Freifunk (Germany), NYC Mesh
  • Advantages: No central ISP, community-controlled
  • Challenges: Limited range, technical complexity

Layer 2 Alternatives: Network Protocols §

Problem: DNS centralization (ICANN controls root servers, domain seizures possible)

Alternatives:

Blockchain-Based Naming:

  • ENS (Ethereum Name Service): Register .eth names on Ethereum blockchain
  • Advantages: Censorship-resistant (no government can seize), permanent ownership
  • Challenges: Requires crypto wallet, expensive (gas fees), not browser-native
  • Namecoin: Bitcoin-based naming system (.bit domains)
  • Similar trade-offs to ENS

Onion Routing:

  • Tor hidden services: Sites with .onion addresses (e.g., 3g2upl4pq6kufc4m.onion for DuckDuckGo)
  • Advantages: Censorship-resistant, anonymous hosting
  • Challenges: Slow, not user-friendly (long random addresses), stigma

Content-Addressed Networking:

  • IPFS: Content identified by hash, not location
  • Example: ipfs://QmXoypizjW3WknFiJnKLwHCnL72vedxjQkDDP1mXWo6uco (instead of http://example.com)
  • Advantages: Content can’t be censored by taking down a domain
  • Challenges: Not human-readable, requires IPFS-enabled browser

Layer 3 Alternatives: Identity Systems §

Problem: Platform-controlled identity (Facebook, Google can delete your account)

Alternatives:

Federated Identity:

  • ActivityPub: @[email protected] (like email addresses)
  • IndieAuth: Use your own domain as identity
  • Advantages: Portable (can migrate), not owned by one company
  • Challenges: Requires domain ownership, instances can still ban you

Decentralized Identifiers (DIDs):

  • Cryptographic identities (public/private key pairs)
  • Not tied to any platform or domain
  • Advantages: Truly sovereign (you control the keys)
  • Challenges: Hard to remember (long strings of characters), key management risky (lose key = lose identity)

Self-Hosted Email:

  • Run your own email server ([email protected])
  • Advantages: Full control, federated (can email anyone)
  • Challenges: Technical (requires sysadmin skills), anti-spam filters often block self-hosted email

Layer 4 Alternatives: Data Storage §

Problem: Cloud lock-in (Google, Dropbox, iCloud can delete your data or change terms)

Alternatives:

Self-Hosted Storage:

  • Nextcloud: Open-source cloud (like Dropbox, but you run it)
  • Syncthing: Peer-to-peer file sync (no central server)
  • Advantages: Full control, no surveillance
  • Challenges: Requires server management, backups are your responsibility

Distributed Storage:

  • IPFS: Content stored across network, no single server
  • Filecoin: Pay for distributed storage with cryptocurrency
  • Arweave: “Permanent” storage (pay once, store forever)
  • Advantages: Censorship-resistant, redundant
  • Challenges: Cost, performance, complexity

Local-First Software:

  • Apps that store data on your device, sync peer-to-peer
  • Examples: Obsidian (notes), Logseq (knowledge management)
  • Advantages: Your data, no cloud dependency
  • Challenges: Syncing across devices is harder

Layer 5 Alternatives: Application Layer §

Problem: Proprietary apps (Twitter, Facebook control how you use their platforms)

Alternatives:

Open Source Apps:

  • Mastodon: Open-source Twitter alternative
  • Pixelfed: Open-source Instagram alternative
  • PeerTube: Open-source YouTube alternative
  • Advantages: Can fork if company sells out, community-governed
  • Challenges: Often lag in features, smaller user bases

Protocol-Based Tools:

  • Email: Any client can access any server
  • RSS: Any reader can subscribe to any feed
  • ActivityPub: Any app can communicate with any instance
  • Advantages: No vendor lock-in, interoperability
  • Challenges: Less shiny than proprietary apps, requires coordination

Layer 6 Alternatives: Economic Layer §

Problem: Payment processor oligopoly (Visa/PayPal can deplatform you)

Alternatives:

Cryptocurrency:

  • Bitcoin, Ethereum, etc.: Censorship-resistant payments
  • Advantages: No bank or processor can block you
  • Challenges: Volatile, technical barriers, environmental concerns, regulatory uncertainty

Cooperatively-Owned Payment Systems:

  • Credit unions: Member-owned banks
  • Payment cooperatives: Users own the payment network
  • Advantages: Democratic control, aligned incentives
  • Challenges: Smaller networks, less convenient

Direct Transactions:

  • Wire transfers, checks, cash: No intermediary
  • Advantages: Can’t be deplatformed
  • Challenges: Slow, expensive, not scalable online

Platform Cooperatives:

  • Stocksy: Photographer-owned stock photo agency
  • Resonate: Musician-owned streaming service
  • Advantages: Workers own the platform, keep more revenue
  • Challenges: Hard to scale, less capital for growth

Part IV: The Sovereignty Stack as Design Framework §

When building new tools, platforms, or institutions, use the Sovereignty Stack as a design checklist:

Sovereignty Audit Template §

For each layer, ask:

Layer 1: Physical Infrastructure

  • [ ] Where are the servers? Who owns them?
  • [ ] What happens if the hosting provider kicks us off?
  • [ ] Do we have redundancy (multiple data centers)?

Layer 2: Network Protocols

  • [ ] Are we using open protocols or proprietary ones?
  • [ ] Can we be censored via DNS (domain seizure)?
  • [ ] Do we have fallback addresses (onion, IPFS, etc.)?

Layer 3: Identity

  • [ ] Do users own their identities (domain-based, cryptographic)?
  • [ ] Can we revoke identities (if so, under what conditions)?
  • [ ] Can users migrate identities to other systems?

Layer 4: Storage

  • [ ] Do users own their data legally?
  • [ ] Can users export everything easily?
  • [ ] Is data stored centrally (vulnerable) or distributed?

Layer 5: Application

  • [ ] Is the software open source (can users fork it)?
  • [ ] Does it depend on our servers, or is it peer-to-peer?
  • [ ] Can other apps interoperate with ours (open protocols)?

Layer 6: Economic

  • [ ] How do we make money without surveillance/extraction?
  • [ ] What payment systems do we use (vulnerable to deplatforming)?
  • [ ] Do we have multiple revenue streams?

Example Audit: Ghost vs. Medium §

Ghost (High Sovereignty):

Layer Assessment
Physical Self-hosted option (users run their own servers) OR Ghost(Pro) hosting (can migrate away)
Network Custom domains supported (users own their URLs)
Identity Domain-based ([email protected])
Storage Users own data, full export tools
Application Open source (can fork if Ghost Ltd dies)
Economic Subscriptions (users pay Ghost, or self-host for free), no ads

Medium (Low Sovereignty):

Layer Assessment
Physical Medium’s servers only
Network medium.com/@username (don’t own domain)
Identity Platform-controlled (Medium can ban you)
Storage Data on Medium’s servers (export exists but limited)
Application Proprietary (can’t fork, can’t self-host)
Economic Ad/subscription-based (users don’t control revenue model)

Result: Ghost embodies more sovereignty across all layers.


Part V: The Political Economy Question — Who Should Own Infrastructure? §

Three Models of Ownership §

Model 1: Corporate Ownership (Status Quo)

  • Private companies own most infrastructure
  • Examples: AWS, Cloudflare, Visa, Google
  • Advantages: Convenient, well-funded, feature-rich
  • Disadvantages: Extractive, surveillance-based, can deplatform users

Model 2: State Ownership

  • Governments own critical infrastructure
  • Examples: Postal service, public utilities, national postal banks
  • Advantages: Democratic accountability (in theory), public service mandate
  • Disadvantages: Vulnerable to authoritarianism, bureaucratic, underfunded

Model 3: Commons Ownership

  • Users collectively own infrastructure
  • Examples: Wikipedia, cooperatives, federated networks
  • Advantages: Aligned incentives, democratic, mission-driven
  • Disadvantages: Coordination challenges, underfunded, slower innovation

Archaeobytology’s Position: Pluralism §

We need all three models for different layers and contexts:

  • Layer 1 (Physical): Mix of corporate, state, and community-owned (diversity prevents single point of failure)
  • Layer 2 (Network): Open protocols (state/community standards-setting, anyone can implement)
  • Layer 3 (Identity): Individual ownership (federated or cryptographic)
  • Layer 4 (Storage): Mix of self-hosted and cooperatively-managed
  • Layer 5 (Application): Open source (community-owned code)
  • Layer 6 (Economic): Multiple options (crypto, co-ops, traditional payment—let users choose)

No single model is perfect. The goal is pluralism: multiple ownership structures coexisting, so users aren’t dependent on any one.


Part VI: Threat Modeling — Attacks on Each Layer §

Threat 1: State Surveillance (All Layers) §

Attack Vector: Governments compel infrastructure providers to surveil users

Examples:

  • NSA tapping undersea cables (Layer 1)
  • Forcing DNS providers to log queries (Layer 2)
  • Demanding user identity records (Layer 3)
  • Accessing cloud storage without warrants (Layer 4)
  • Installing backdoors in apps (Layer 5)
  • Tracking financial transactions (Layer 6)

Defenses:

  • End-to-end encryption (can’t surveil what you can’t read)
  • Jurisdiction shopping (host in privacy-respecting countries)
  • Onion routing (obscure who’s talking to whom)
  • Decentralization (harder to compel many actors than one)

Threat 2: Corporate Extraction (Layers 4-6) §

Attack Vector: Companies monetize user data and attention

Examples:

  • Surveillance advertising (Layer 4: data mining)
  • Algorithmic manipulation (Layer 5: addictive app design)
  • Platform rent (Layer 6: taking 30% cuts)

Defenses:

  • Business models that don’t require surveillance (subscriptions, donations)
  • Open source alternatives (no company owns the app)
  • Data ownership (users control, export, delete)

Threat 3: Platform Capture (Layers 3-5) §

Attack Vector: Dominant platforms use network effects to lock in users

Examples:

  • Identity capture (can’t leave Facebook without losing social graph)
  • Data lock-in (hard to export and migrate)
  • API restrictions (third-party apps banned)

Defenses:

  • Interoperability (force platforms to allow data portability)
  • Federation (multiple providers, users can migrate)
  • Open protocols (anyone can build competing apps)

Threat 4: Infrastructure Fragility (Layer 1) §

Attack Vector: Physical infrastructure fails or is attacked

Examples:

  • Undersea cable cuts (accidental or sabotage)
  • Data center fires (OVH fire in 2021 destroyed thousands of sites)
  • Power grid failures (take down entire regions)

Defenses:

  • Redundancy (multiple data centers, geographic distribution)
  • Peer-to-peer architectures (no single server to attack)
  • Regular backups (distributed across locations)

Conclusion: Sovereignty Requires Stack Thinking §

You cannot achieve digital sovereignty by fixing one layer. Even if you:

  • Own your domain (Layer 2)
  • Use open source software (Layer 5)
  • Self-host your data (Layer 4)

…you’re still vulnerable if:

  • Your hosting provider kicks you off (Layer 1)
  • Payment processors deplatform you (Layer 6)
  • Your identity depends on a platform (Layer 3)

Sovereignty requires thinking across all six layers. It’s not enough to own one part of the stack; you need resilience at every level.

This is hard. Full-stack sovereignty is expensive, technical, and time-consuming. Most people will accept some vulnerability in exchange for convenience. That’s okay—sovereignty is a spectrum, not binary.

But the Sovereignty Stack helps you:

  • Diagnose vulnerabilities: Where are you most at risk?
  • Prioritize fixes: Which layer matters most for your use case?
  • Design resilient systems: How do you build across all layers?
  • Evaluate tools: Does this app/platform embody sovereignty?

In the next chapter, we’ll explore movement building—how to turn individual sovereignty into collective power. Because infrastructure isn’t just technical; it’s political. Changing who controls the stack requires organizing, advocacy, and policy change.

The political economy of Ground isn’t just about building alternatives. It’s about fighting for a different kind of internet—one where users, not corporations or states, have power.

That fight begins with understanding the Stack. Now you do.


Discussion Questions §

  1. Personal Sovereignty Audit: Use the Sovereignty Stack to audit your own digital life. Which layers are you most vulnerable on? Which are you most secure on?

  2. Trade-offs: Would you accept less convenience for more sovereignty? Where’s your breaking point? (e.g., Would you run your own email server?)

  3. Threat Prioritization: Which threat is most dangerous: state surveillance, corporate extraction, platform capture, or infrastructure fragility? Does it depend on your context?

  4. Ownership Models: Should critical infrastructure (DNS, payment systems, cloud hosting) be owned by corporations, governments, or commons? What’s best for sovereignty?

  5. Layer Interdependence: If you could only secure ONE layer of the stack, which would you choose? Why?

  6. Future Scenario: Imagine 2040. What does digital infrastructure look like if Archaeobytology succeeds? What if it fails?


Exercise: Redesign One Layer §

Task: Choose one layer of the Sovereignty Stack. Redesign it to maximize sovereignty while remaining practically viable.

Part 1: Critique Current State (500 words)

  • How does the current system work?
  • Who controls it?
  • What are the sovereignty failures?
  • What specific harms result from current design?

Part 2: Design Alternative (1000 words)

  • How would you redesign this layer?
  • What technologies enable your design?
  • How does it resist capture (corporate, state, platform)?
  • How does it balance sovereignty with usability?

Part 3: Adoption Strategy (500 words)

  • How do you get people to switch?
  • What’s the transition path from current system to yours?
  • What are the obstacles (technical, economic, political)?

Part 4: Threat Model (500 words)

  • How would each threat actor attack your system?
  • State surveillance
  • Corporate extraction
  • Platform capture
  • Infrastructure failure
  • What defenses have you built in?
  • What vulnerabilities remain?

Part 5: Reflect (300 words)

  • What surprised you about this design exercise?
  • What compromises did you have to make?
  • Would you personally use the system you designed?

Further Reading §

On Infrastructure and Power §

  • Lessig, Lawrence. Code: Version 2.0. Basic Books, 2006.
  • Classic on how digital architecture encodes power

  • Star, Susan Leigh. “The Ethnography of Infrastructure.” American Behavioral Scientist 43, no. 3 (1999): 377-391.

  • Infrastructure is political, not neutral

  • Winner, Langdon. “Do Artifacts Have Politics?” Daedalus 109, no. 1 (1980): 121-136.

  • Technologies embody political choices

On Specific Layers §

  • Mueller, Milton. Networks and States: The Global Politics of Internet Governance. MIT Press, 2010.
  • On DNS, ICANN, and network governance

  • Schneier, Bruce. Data and Goliath. W.W. Norton, 2015.

  • On surveillance at every layer

  • DeNardis, Laura. The Global War for Internet Governance. Yale, 2014.

  • Who controls protocols and standards

On Alternatives §

  • Doctorow, Cory. The Internet Con: How to Seize the Means of Computation. Verso, 2023.
  • On interoperability and adversarial compat

  • Bauwens, Michel, and Vasilis Kostakis. Network Society and Future Scenarios for a Collaborative Economy. Palgrave, 2014.

  • On commons-based alternatives

  • Schneider, Nathan. “Cryptoeconomics as a Limitation on Governance.” COALA Workshop, 2017.

  • On blockchain governance and limitations

Primary Sources §

  • Tor Project. “How Tor Works.” https://www.torproject.org/about/history/
  • IPFS Documentation. https://docs.ipfs.tech/
  • ENS Documentation. https://docs.ens.domains/
  • Mastodon Documentation. https://docs.joinmastodon.org/

End of Chapter 15

Next: Chapter 16 — From Practice to Discipline: Movement Building

Part IV • Disciplinary Movement & The Post-Platform Future

Chapter 16: From Practice to Discipline

Movement Building

24 min read 5,168 words

Opening: The Marathon Nobody Knows You’re Running §

In 1949, a small group of scholars gathered at MIT to discuss “the possibilities of a science of science.” They called themselves historians and sociologists of science, though neither history departments nor sociology departments particularly wanted them. They were too historical for sociologists, too sociological for historians, too focused on content for both.

By 1975, they’d founded the Society for Social Studies of Science (4S). By 1990, there were doctoral programs at MIT, Cornell, and Edinburgh. By 2000, Science and Technology Studies (STS) was recognized as a legitimate interdisciplinary field with journals, conferences, and tenure-track jobs.

It took 50 years.

In 2012, a data scientist named DJ Patil coined the term “data science” (building on earlier uses). Tech companies were desperate for people who could analyze big data but didn’t know what to call them. Universities scrambled to create programs. By 2020, data science was everywhere—hundreds of degree programs, professional certifications, six-figure salaries.

It took 8 years.

One discipline took half a century to build through patient coalition-building, scholarly legitimation, and institutional negotiation. The other exploded in less than a decade driven by industry demand and money.

Archaeobytology faces the same question every emerging discipline does: How do we go from scattered practice to recognized field?

Do we take the slow road—building scholarly infrastructure, publishing rigorous research, waiting for academic legitimacy? Or the fast road—chasing industry funding, training practitioners, proving economic value?

The answer is: both, strategically, over 10-20 years.

This chapter is your roadmap. By the end, you’ll understand:

  • The five dimensions of movement building (knowledge, institutions, careers, visibility, policy)
  • Case studies of successful discipline formation (DH, Data Science, STS)
  • A phased timeline for Archaeobytology (what to build when)
  • How to avoid common failure modes (capture, fragmentation, irrelevance)
  • What YOU can do right now to help

This isn’t just theory. This is praxis—the strategic work of turning an idea into institutional reality.

Let’s begin.


Part I: The Movement-Building Matrix §

The Five Dimensions §

Every successful discipline requires infrastructure across five dimensions. Neglect any one, and the movement stalls.

Dimension 1: Knowledge Infrastructure

What it is: The intellectual scaffolding that makes a field coherent.

Components:

  • Journals: Peer-reviewed venues for publishing research
  • Conferences: Annual gatherings to share work and build community
  • Textbooks: Standardized curriculum (this book is one example)
  • Handbooks: Reference works covering methods, theory, history
  • Online platforms: Wikis, forums, repositories for distributed knowledge

Why it matters: Without knowledge infrastructure, practitioners can’t:

  • Find each other’s work (no central publication venue)
  • Build on prior research (no shared literature)
  • Train students (no textbooks, no canon)
  • Claim intellectual coherence (no shared vocabulary)

Archaeobytology’s Current State (2025):

  • ✅ This textbook exists
  • ❌ No dedicated journal (yet)
  • ❌ No annual conference (yet)
  • ❌ No handbook (yet)
  • ⚠️ Scattered online communities (Archive Team wiki, IndieWeb, but no unified Archaeobytology hub)

Priority Actions:

  1. Launch Journal of Archaeobytology (open access, online)
  2. Host first Archaeobytology conference (even if small—50 people)
  3. Create archaeobytology.org wiki (methods, case studies, tools)

Dimension 2: Institutional Anchors

What it is: Physical/organizational homes where the discipline can grow.

Components:

  • Departments: Standalone units with hiring/budgeting autonomy
  • Programs: Degree-granting (certificates, minors, majors, graduate programs)
  • Centers/Institutes: Research hubs (may not grant degrees but provide infrastructure)
  • Labs: Spaces with equipment, servers, staff
  • Professional schools: Practice-oriented training (like law schools, business schools)

Why it matters: Without institutional anchors, the field is:

  • Homeless: No physical space, no servers, no resources
  • Jobless: No tenure-track positions for PhDs
  • Powerless: No budgets, no hiring authority, no institutional clout

Archaeobytology’s Current State (2025):

  • ❌ No standalone departments
  • ❌ No degree programs (though scattered courses exist)
  • ⚠️ Internet Archive functions as de facto institute (but not academic)
  • ❌ No university-based centers (yet)

Priority Actions:

  1. Launch “Certificate in Digital Preservation and Sovereignty” at 3-5 universities
  2. Establish “Center for Archaeobytology” at one major university (with grant funding)
  3. Create first MA program (likely in iSchool or interdisciplinary program)

Dimension 3: Professional Pathways

What it is: Jobs people can get after training in the field.

Tracks:

  • Academic: Tenure-track faculty, postdocs, research positions
  • Practitioner: Archivists, curators, preservation specialists in libraries/museums
  • Industry: Tech companies (digital sovereignty engineers, ethical AI trainers)
  • Non-profit: Internet Archive, EFF, Creative Commons, Wikimedia roles
  • Consulting: Freelance/agency work advising on preservation and platform alternatives
  • Government: National Archives, Library of Congress, policy roles

Why it matters: Students won’t enroll in programs if there are no jobs. Universities won’t create programs if they can’t place graduates.

Archaeobytology’s Current State (2025):

  • ⚠️ Some relevant jobs exist (digital archivist, preservation specialist) but don’t use “Archaeobytology” term
  • ❌ No clear career ladder (junior → senior → leadership)
  • ❌ No professional certification (no “Certified Archaeobytologist” credential)

Priority Actions:

  1. Survey existing jobs and map to Archaeobytology skills
  2. Create “Certified Archaeobytologist” credential (like Certified Archivist)
  3. Build job board (archaeobytology.org/jobs)
  4. Develop clear career pathways document (“If you get an MA in Archaeobytology, you can work as…”)

Dimension 4: Public Visibility

What it is: Awareness outside academia—general public, media, policymakers.

Mechanisms:

  • Popular books: Trade press (not just academic presses)
  • Podcasts: Storytelling and interviews
  • Documentaries: Visual media for mass audiences
  • Op-eds: New York Times, Atlantic, Wired, etc.
  • Social media: Twitter/Mastodon accounts, YouTube channels
  • TED talks: High-profile speaking (reaches millions)
  • Museum exhibits: Physical installations about platform death

Why it matters: Academic legitimacy alone isn’t enough. Public visibility:

  • Attracts students (people major in things they’ve heard of)
  • Influences funders (foundations fund visible causes)
  • Shapes policy (legislators care about issues the public cares about)
  • Creates urgency (media coverage makes platform death a “real” problem)

Archaeobytology’s Current State (2025):

  • ⚠️ Some media coverage of platform shutdowns (but not framed as “Archaeobytology”)
  • ⚠️ Internet Archive gets press, but not as discipline-building
  • ❌ No breakout popular book (need the “gladwell moment”)
  • ❌ No documentary

Priority Actions:

  1. Write popular book on platform death (trade press, accessible prose)
  2. Produce documentary: “The Day GeoCities Died” or “Who Killed Your Childhood Website?”
  3. Get 5-10 op-eds in major outlets
  4. Launch public-facing podcast: “Murdered Platforms” (each episode covers one shutdown)

Dimension 5: Policy Advocacy

What it is: Translating research into laws, regulations, and norms.

Policy Goals:

  • Right to Archive: Expand fair use/copyright exceptions for preservation
  • Platform Accountability: Require notice before shutdowns, mandate data export tools
  • Digital Right of First Refusal: Archives get access to content before deletion
  • Public Digital Archive Funding: Dedicated government funding stream (like NEH but for digital preservation)
  • Anti-Speculation Measures: Prevent domain squatting, ensure use-it-or-lose-it for digital infrastructure

Mechanisms:

  • White papers: Research reports with policy recommendations
  • Testimony: Speaking at legislative hearings
  • Model legislation: Draft bills ready for lawmakers to introduce
  • Coalition building: Partner with EFF, Internet Archive, library associations, tech policy orgs
  • Informal briefings: Meet with congressional staffers, regulators

Why it matters: Scholarly work alone doesn’t change systems. Laws shape:

  • What can be archived legally
  • Whether platforms must provide data export tools
  • Whether governments fund preservation infrastructure
  • Whether digital culture is protected like tangible cultural heritage

Archaeobytology’s Current State (2025):

  • ⚠️ Some advocacy happening (Internet Archive lawsuits, EFF campaigns) but not framed as Archaeobytology movement
  • ❌ No unified policy agenda
  • ❌ No Archaeobytologist testifying at hearings (yet)

Priority Actions:

  1. Draft “Archaeobytologist’s Policy Agenda” (5-10 key legislative goals)
  2. Form “Coalition for Digital Preservation Rights” (partner orgs)
  3. Get first Archaeobytologist to testify at congressional hearing
  4. Publish white paper: “The Case for a Right to Archive”

Part II: Case Studies in Discipline Formation §

Case Study 1: Digital Humanities (40-Year Marathon) §

Timeline:

1960s-1980s: Scattered Practice

  • “Humanities computing”—scholars using computers for text analysis
  • No community, no infrastructure, seen as technical skill not intellectual field

1990s: Early Organization

  • Conferences emerge: ACH (1978 but small), TEI (Text Encoding Initiative, 1987)
  • First journals: Computers and the Humanities (1966, but niche)
  • Internet makes digital methods suddenly relevant

2000s: Critical Mass

  • Term “digital humanities” replaces “humanities computing” (2004)
  • Major conferences: DH (annual, hundreds of attendees)
  • Centers at Stanford, UVA, CUNY, Nebraska
  • NEH Office of Digital Humanities (2008)—dedicated funding

2010s: Institutionalization

  • Dozens of DH centers worldwide
  • Hundreds of tenure-track jobs with “DH” in title
  • Textbooks, handbooks, journals proliferate
  • Still fights for legitimacy (but no longer dismissed as “not real scholarship”)

2020s: Established but Marginal

  • DH is recognized field
  • Still mostly interdisciplinary (few standalone departments)
  • Debates about boundaries, politics, labor conditions

Key Lessons:

Slow and steady wins legitimacy—took 40 years but built durable infrastructure

External funding helps—NEH Office of Digital Humanities accelerated growth

Rebranding matters—“digital humanities” sounded more intellectual than “humanities computing”

Still marginal—even after 40 years, many DH scholars struggle for tenure

Labor exploitation—lots of adjuncts/alt-ac, few permanent positions

For Archaeobytology:

  • Don’t expect fast legitimation (but aim for faster than 40 years)
  • Pursue NEH/Mellon funding aggressively
  • Name matters (Archaeobytology is provocative, good)
  • Build labor protections from start (don’t replicate DH’s precarity)

Case Study 2: Data Science (Industry-Driven Speedrun) §

Timeline:

2000s: Industry Need

  • Companies drowning in data, no one trained to analyze it
  • Hiring statisticians, CS PhDs, physicists—anyone who could code + math

2008-2012: Term Emerges

  • DJ Patil and Jeff Hammerbacher coin “data scientist” (2008-2012)
  • Industry demand explodes (Google, Facebook, Amazon hiring aggressively)

2012-2015: Academic Response

  • Universities see $$$ (lucrative master’s programs)
  • Bootcamps emerge (Galvanize, General Assembly—3-6 month training)
  • Columbia, NYU, UC Berkeley launch data science programs

2015-2020: Ubiquity

  • Hundreds of programs (undergrad, master’s, PhD)
  • Data Science Society, professional certifications
  • Six-figure salaries attract students

2020s: Established but Fuzzy

  • Data science everywhere
  • Still debates about what it “is” (statistics? CS? business analytics?)
  • Academic programs vary wildly in quality

Key Lessons:

Industry demand accelerates everything—8 years to ubiquity

Money talks—universities created programs because students would pay

Bootcamps work—don’t need PhD to be data scientist, practical training suffices

Intellectual incoherence—field still doesn’t have clear boundaries or canon

Quality control—some programs are excellent, many are cash grabs

For Archaeobytology:

  • Identify industry demand (tech companies need digital sovereignty architects?)
  • Create “Archaeobytology Bootcamp” (3-6 month intensive, professional credential)
  • But maintain intellectual rigor (don’t let money corrupt mission)
  • Build quality standards early

Case Study 3: Science and Technology Studies (Coalition Model) §

Timeline:

1970s: Coalition Formation

  • Historians of science + sociologists of knowledge + philosophers of technology
  • All studying science/tech but from different angles
  • Realized they had shared interests → formed coalition

1975: Professional Society

  • Founded 4S (Society for Social Studies of Science)
  • Annual conference becomes gathering place

1980s-1990s: Boundary Struggles

  • “Science Wars”—scientists attack STS as postmodern relativism
  • Internal debates: constructivism vs. realism, Latour vs. feminists
  • Field almost fractures but holds together

2000s: Stabilization

  • Multiple journals (Social Studies of Science, Science, Technology & Human Values)
  • Departments at MIT, Cornell, UC San Diego, York, others
  • Clear identity: interdisciplinary but distinct

2010s-2020s: Maturity

  • Hundreds of STS scholars worldwide
  • Influencing policy (COVID response, climate, AI ethics)
  • Still interdisciplinary (mostly joint appointments) but recognized

Key Lessons:

Coalitions work—united historians, sociologists, philosophers under one tent

Boundary struggles are normal—every field fights over what it is/isn’t

Professional society matters—4S gave STS institutional home

Interdisciplinarity can be strength—not having disciplinary “purity” allows flexibility

Slow growth—40+ years, still mostly joint appointments not standalone departments

For Archaeobytology:

  • Build coalition across digital historians, archivists, activists, builders
  • Expect internal debates (Archive vs Anvil priorities, etc.)—that’s healthy
  • Found professional society early (within 5 years)
  • Embrace interdisciplinarity as strength

Part III: The Archaeobytology Movement Strategy (10-20 Year Roadmap) §

Phase 1: Emergence (Years 1-5) — WE ARE HERE §

Current State (2025):

  • Scattered practitioners doing Archaeobytology without calling it that
  • This textbook is one of first attempts to codify field
  • No formal infrastructure (yet)
  • ~50-100 people might identify as doing this work (but don’t use “Archaeobytology” term)

Goals for Years 1-5:

Year 1 (2025-2026):

  • [ ] Publish this textbook (archaeobytology.org)
  • [ ] Create wiki: archaeobytology.org/wiki (methods, case studies, tools)
  • [ ] Launch mailing list/Discord for practitioners
  • [ ] Collect 100 email addresses of people interested
  • [ ] Host first “Archaeobytology Unconference” (virtual, 1 day, informal)

Year 2 (2026-2027):

  • [ ] Launch Journal of Archaeobytology (open access, online-only at first)
  • [ ] Solicit 10 papers for inaugural issue
  • [ ] Host first in-person conference: “Archaeobytology 2027” (50-100 people)
  • [ ] Secure first grant (Mellon/NEH/Mozilla for $50-100k)
  • [ ] 5 universities offer “Introduction to Archaeobytology” course

Year 3 (2027-2028):

  • [ ] Second annual conference (100-150 people)
  • [ ] Publish second journal issue (aim for 2/year)
  • [ ] Launch first certificate program (one university offers “Certificate in Digital Preservation”)
  • [ ] Draft “Archaeobytologist’s Policy Agenda” white paper
  • [ ] 10 universities teaching Archaeobytology courses

Year 4 (2028-2029):

  • [ ] Found “Society for Archaeobytology” (or similar name)
  • [ ] Third conference (150-200 people)
  • [ ] First student graduates with certificate in Archaeobytology (milestone!)
  • [ ] Publish popular article in Atlantic or Wired
  • [ ] Secure larger grant ($200-500k) for multi-year project

Year 5 (2029-2030):

  • [ ] Journal has 4 issues/year, editorial board of 20
  • [ ] Conference has 250+ attendees, multiple tracks
  • [ ] 3 universities have certificates/minors
  • [ ] First testimony at Congressional hearing by Archaeobytologist
  • [ ] 20+ universities teaching courses

Phase 1 Success Metrics:

  • ✅ 500+ people identify as Archaeobytologists
  • ✅ Professional society exists
  • ✅ Annual conference established
  • ✅ Journal publishing regularly
  • ✅ Some undergraduate programs

Phase 2: Coalition Building (Years 6-10) §

Goals for Years 6-10:

Infrastructure:

  • [ ] Journal becomes quarterly, peer-reviewed, indexed (Scopus, Web of Science)
  • [ ] Conference grows to 500 attendees, international
  • [ ] Handbook published: Handbook of Archaeobytology (40+ chapters, major reference work)
  • [ ] Online platform mature (wiki has 1000+ pages, forum has 5000+ members)

Institutions:

  • [ ] First MA program launches (probably iSchool or interdisciplinary)
  • [ ] 3-5 universities have “Centers for Digital Sovereignty” (funded, with staff)
  • [ ] 10+ universities have certificate/minor programs
  • [ ] First PhD student lists “Archaeobytology” as primary field (even if in interdisciplinary program)

Careers:

  • [ ] 50+ tenure-track jobs posted with “Archaeobytology” or “Digital Sovereignty” in description
  • [ ] “Certified Archaeobytologist” credential launched (professional certification)
  • [ ] Job placement rate for MA graduates: 80%+

Visibility:

  • [ ] Popular book published (trade press): Murdered Platforms: The Fight for Digital Memory
  • [ ] Documentary released: screening at festivals, streaming on Netflix/Prime
  • [ ] 20+ op-eds in major outlets
  • [ ] Podcast has 50+ episodes, 10k+ listeners

Policy:

  • [ ] “Coalition for Digital Preservation Rights” formed (10+ partner orgs)
  • [ ] Model legislation drafted (“Digital Preservation Act”)
  • [ ] 3+ Archaeobytologists testify at hearings
  • [ ] First local/state policy win (e.g., state library system adopts Archaeobytology standards)

Phase 2 Success Metrics:

  • ✅ 2,000+ Archaeobytologists worldwide
  • ✅ MA programs at 5+ universities
  • ✅ First dissertations completed
  • ✅ Public awareness: 10% of people have heard of Archaeobytology
  • ✅ Policy engagement: regular testimony, coalition work

Phase 3: Institutionalization (Years 11-15) §

Goals for Years 11-15:

Academic Maturity:

  • [ ] 5+ PhD programs offer Archaeobytology as concentration/specialization
  • [ ] 50+ dissertations completed
  • [ ] 100+ tenure-track faculty
  • [ ] Textbook adoption: 100+ universities using this or similar books

Institutional Expansion:

  • [ ] First standalone “Department of Archaeobytology and Digital Sovereignty”
  • [ ] 20+ centers/institutes worldwide
  • [ ] Major research universities (Harvard, MIT, Stanford, etc.) have programs

Funding Ecosystem:

  • [ ] NSF creates “Digital Sovereignty and Preservation” program
  • [ ] NEH has dedicated Archaeobytology funding stream ($5-10M/year)
  • [ ] Private foundations (Mellon, Sloan, Knight) regularly fund Archaeobytology projects

Public Impact:

  • [ ] New York Times runs major feature: “The Archaeobytologists Saving the Internet”
  • [ ] TED Talk by prominent Archaeobytologist (1M+ views)
  • [ ] Museum exhibits at major institutions (Smithsonian, V&A, etc.)

Policy Wins:

  • [ ] Federal legislation passed (e.g., “Digital Preservation Act” or similar)
  • [ ] Library of Congress has “Archaeobytology Division”
  • [ ] International policy (UNESCO recognizes digital cultural heritage, influenced by Archaeobytology work)

Phase 3 Success Metrics:

  • ✅ 5,000+ Archaeobytologists
  • ✅ 100+ universities with programs
  • ✅ First standalone departments
  • ✅ Regular federal funding
  • ✅ Major policy wins

Phase 4: Maturity and Expansion (Years 16-20) §

Goals for Years 16-20:

Discipline Established:

  • [ ] 10+ standalone departments
  • [ ] 1,000+ PhD holders
  • [ ] Archaeobytology included in standard university catalogs (alongside History, Sociology, etc.)
  • [ ] Canon established (everyone agrees on core texts to read)

Global Reach:

  • [ ] Archaeobytology programs in 20+ countries
  • [ ] International professional societies (European Archaeobytology Association, Asia-Pacific chapter, etc.)
  • [ ] Multilingual scholarship (not just English-language dominance)

Specialization:

  • [ ] Subfields emerge: “Archaeobytology of Social Media,” “Video Game Preservation Studies,” “Digital Memory and Trauma,” etc.
  • [ ] Specialized journals for subfields
  • [ ] Debates about what “counts” as Archaeobytology (sign of maturity)

Cultural Impact:

  • [ ] High school students learn about platform death in history classes
  • [ ] Archaeobytology consultants common (like “sustainability consultants” today)
  • [ ] Major corporations hire Archaeobytologists (ethical concerns, but shows mainstream acceptance)

Phase 4 Success Metrics:

  • ✅ 10,000+ Archaeobytologists worldwide
  • ✅ Field is recognized and established
  • ✅ Career pathways clear and diverse
  • ✅ Public awareness: majority of people have heard of Archaeobytology

Part IV: Avoiding Common Failure Modes §

Failure Mode 1: Disciplinary Capture §

Risk: Existing fields absorb Archaeobytology, prevent independence.

Scenario:

  • History departments say: “Archaeobytology is just digital history, we’ll hire one person for that”
  • CS departments say: “We’ll add a preservation course, that’s enough”
  • Library schools say: “Web archiving covers this already”
  • Result: Archaeobytology gets fragmented, never achieves critical mass

Defense:

  • Insist on synthesis: Archaeobytology is NOT reducible to any single field
  • Build independent infrastructure: Our own journals, conferences, society (harder to absorb)
  • Coalition across departments: If History, CS, AND Library Science all want us, harder for any one to capture
  • Interdisciplinary programs: Don’t let traditional departments be gatekeepers

Failure Mode 2: Industry Co-optation §

Risk: Tech companies use Archaeobytology rhetoric but corrupt mission.

Scenario:

  • Facebook hires “Digital Preservation Specialists” (to preserve user data for ad targeting, not user sovereignty)
  • Blockchain startups claim to be “Archaeobytological” (conflating crypto speculation with preservation)
  • Archaeobytology jobs become corporate compliance roles (ethics-washing)
  • Result: Field becomes associated with surveillance capitalism, loses critical edge

Defense:

  • Value clarity: Center the Three Pillars in everything (Declaration, Connection, Ground)
  • Ethical guidelines: Professional code that says “surveillance-capitalism work is not Archaeobytology”
  • Critical scholarship: Maintain academic independence, publish critiques of platform power
  • Diversity of employment: Balance academic, non-profit, and (selective) industry roles

Failure Mode 3: Internal Fragmentation §

Risk: Practitioners can’t agree on boundaries, methods, values → field splinters.

Scenario:

  • “Archive-first” Archaeobytologists vs. “Anvil-first” Archaeobytologists fight
  • “Everything should be preserved” camp vs. “Consent is paramount” camp can’t reconcile
  • Methodological wars: “Only bit-perfect forensics count” vs. “Triage means good-enough”
  • Result: No unified identity, people stop using “Archaeobytology” label, movement dissolves

Defense:

  • Big tent philosophy: Multiple approaches valid, don’t excommunicate over disagreements
  • Core values, flexible methods: Agree on Three Pillars and Custodial Filter, but allow methodological diversity
  • Productive debate: Disagreement is healthy (sign of intellectual vitality), but don’t let it become toxic
  • Generosity: Assume good faith, even when you disagree

Failure Mode 4: Funding Drought §

Risk: Foundations/agencies don’t fund Archaeobytology, infrastructure collapses.

Scenario:

  • Economic recession cuts humanities funding
  • Political shifts defund preservation and digital rights
  • Competing priorities (AI, climate) absorb available grants
  • Result: Journals fold, conferences stop, centers close, people leave for funded fields

Defense:

  • Diversify funding: Don’t depend on one source (get government + foundation + individual donations + earned revenue)
  • Demonstrate impact: Show funders that Archaeobytology matters (saves culture, influences policy, creates jobs)
  • Build endowment: If successful, create financial cushion (like established disciplines have)
  • Partnerships: Work with stable institutions (libraries, museums with guaranteed budgets)

Failure Mode 5: Elitism and Gatekeeping §

Risk: Field becomes exclusive club, shuts out marginalized practitioners.

Scenario:

  • “Real Archaeobytologists” have PhDs from elite universities
  • Practitioners without credentials dismissed (even if doing excellent work)
  • Field replicates academia’s racism, sexism, classism
  • Result: Narrow, homogeneous community that doesn’t reflect diversity of digital culture

Defense:

  • Multiple pathways: PhDs, certificates, self-taught practitioners all valid
  • Open access: Free textbooks, free journals, free conference options
  • Anti-discrimination: Explicit commitments to equity, diverse leadership
  • Value practice: Don’t privilege academic theory over applied work (both matter)
  • Community accountability: Call out gatekeeping when it happens

Part V: What You Can Do Right Now §

If You’re a Student §

Immediate (This Week):

  1. Call yourself an Archaeobytologist—in your bio, on your CV, on social media
  2. Start a reading group—gather 3-5 friends, work through this textbook
  3. Join online communities—find Archive Team, IndieWeb, digital preservation groups

Short-term (This Semester):

  1. Write a paper using Archaeobytology framework—apply Three Pillars, Custodial Filter, etc. to your research
  2. Propose an independent study—pitch “Introduction to Archaeobytology” to sympathetic professor
  3. Start a blog—document your learning, build public portfolio

Medium-term (This Year):

  1. Attend a conference—submit to ADHO, SAA, 4S, or organize Archaeobytology session
  2. Contribute to a project—volunteer with Archive Team, Internet Archive, etc.
  3. Build something—create a tool, preserve a dying platform, start an archive

If You’re a Practitioner §

Immediate:

  1. Document your work—write tutorials, case studies, method posts
  2. Publish—submit to journals, blogs, preprint servers
  3. Teach—offer workshop at local library, hackerspace, or online

Short-term:

  1. Organize a meetup—gather local practitioners, even if just 5 people
  2. Propose conference session—at existing conference, submit “Archaeobytology panel”
  3. Seek funding—apply for grant explicitly for “Archaeobytology research”

Medium-term:

  1. Mentor students—take on interns, advise theses
  2. Build partnerships—connect with libraries, museums, universities
  3. Advocate—write op-ed, contact your representative about digital preservation

If You’re a Professor §

Immediate:

  1. Teach a course—offer “Introduction to Archaeobytology” (use this textbook)
  2. Cite Archaeobytology—in your research, explicitly name the field
  3. Advise students—encourage dissertations in Archaeobytology

Short-term:

  1. Organize working group—gather colleagues across departments interested in this work
  2. Apply for grant—propose “Center for Digital Sovereignty” or similar
  3. Hire—when job openings come, advocate for Archaeobytology specialization

Medium-term:

  1. Create program—certificate, minor, or master’s in Archaeobytology
  2. Launch journal—start Journal of Archaeobytology at your university press
  3. Host conference—organize first major Archaeobytology conference at your institution

If You’re an Administrator §

Immediate:

  1. Support faculty—when they propose Archaeobytology courses/programs, approve them
  2. Fund infrastructure—allocate space, servers, staff support
  3. Strategic hire—create position in Archaeobytology (signal to field it’s legitimate)

Short-term:

  1. Create certificate program—low-cost way to test demand
  2. Partner with institutions—connect with Internet Archive, local libraries
  3. Seek external funding—apply for grants to create center/program

Medium-term:

  1. Launch degree program—MA in Archaeobytology (draws students, generates revenue)
  2. Build center—dedicate space and staff to Archaeobytology research/teaching
  3. Advocate—tell peer institutions, accreditors, funders that this field matters

Part VI: Movement Coalitions and Alliances §

Who Are Our Natural Allies? §

1. Librarians and Archivists

  • Shared interests: Preservation, access, metadata, long-term stewardship
  • Partnerships: Joint programs, shared infrastructure, professional development
  • Organizations: SAA (Society of American Archivists), ALA (American Library Association)

2. Digital Humanists

  • Shared interests: Digital methods, scholarly infrastructure, interdisciplinarity
  • Partnerships: Joint conferences, share faculty lines, collaborative research
  • Organizations: ADHO (Alliance of Digital Humanities Organizations)

3. Digital Rights Activists

  • Shared interests: Platform accountability, user sovereignty, right to archive
  • Partnerships: Policy advocacy, public campaigns, legal challenges
  • Organizations: EFF (Electronic Frontier Foundation), Creative Commons, Internet Archive

4. Tech Workers and Ethical Engineers

  • Shared interests: Building alternatives, open protocols, resistance to surveillance capitalism
  • Partnerships: Tool-building, technical consulting, job placements
  • Organizations: Tech Workers Coalition, Worker cooperatives

5. STS Scholars

  • Shared interests: Studying platform power, technological politics, social construction of technology
  • Partnerships: Theoretical frameworks, joint research, publishing
  • Organizations: 4S (Society for Social Studies of Science)

6. Museums and Memory Institutions

  • Shared interests: Interpreting artifacts, public engagement, cultural heritage
  • Partnerships: Exhibitions, public programs, institutional preservation
  • Organizations: ICOM (International Council of Museums), AAM (American Alliance of Museums)

Building the Coalition §

Strategy 1: Multi-Stakeholder Convenings

  • Host annual “Digital Preservation Summit” bringing together all allied groups
  • Not just Archaeobytologists—invite librarians, activists, engineers, scholars, policymakers
  • Goal: Build shared agenda while respecting different priorities

Strategy 2: Cross-Organizational Membership

  • Encourage Archaeobytologists to join SAA, ADHO, 4S, EFF
  • Present at their conferences, publish in their journals
  • Don’t isolate—embed ourselves in adjacent communities

Strategy 3: Shared Infrastructure

  • Offer to host Archaeobytology track at existing conferences (before we have our own)
  • Publish in existing journals (while also building our own)
  • Use existing organizations’ resources (mailing lists, platforms) early on

Strategy 4: Policy Coalitions

  • Form “Alliance for Digital Preservation Rights” (umbrella org)
  • Members: Archaeobytologists, libraries, Internet Archive, EFF, academics, tech workers
  • Unified policy agenda: Right to Archive, Platform Accountability, Public Funding

Conclusion: The Long Game §

Building a discipline takes patience, strategy, and collective will.

Digital Humanities took 40 years. Data Science took 8 (but with massive industry backing). STS took 40 (but created durable coalitions).

Archaeobytology’s timeline: Somewhere in between. With strategic action, we could achieve:

  • 5 years: Professional society, annual conference, first certificates
  • 10 years: MA programs, regular funding, public visibility
  • 15 years: PhD programs, departments, policy influence
  • 20 years: Fully established discipline

This won’t happen automatically. It requires:

  • Students declaring “I am an Archaeobytologist” (identity formation)
  • Practitioners publishing, teaching, building (knowledge creation)
  • Professors creating programs, hiring, securing grants (institutionalization)
  • Administrators supporting infrastructure (resources)
  • Everyone organizing, advocating, collaborating (movement building)

You are not just reading about a discipline. You are helping build it.

Every time you:

  • Use “Archaeobytology” in your work (you legitimize the term)
  • Cite this textbook (you build canon)
  • Teach a course (you train next generation)
  • Preserve an artifact (you do the work)
  • Advocate for policy (you change systems)
  • Mentor a student (you grow the field)

…you are building the movement.

In 20 years, there might be Archaeobytology departments at universities. Students might major in it. Laws might protect digital culture because we advocated for them.

Or not. That depends on us.

The marathon has begun. You’re running it whether you know it or not.

Now: Run intentionally. Run together. Run toward the finish line.

The discipline we need is the discipline we build.


Discussion Questions §

  1. Personal Role: In the Movement-Building Matrix (knowledge, institutions, careers, visibility, policy), which dimension are you best positioned to contribute to? Why?

  2. Timeline Realism: Is a 20-year timeline realistic? Too optimistic? Too pessimistic? What would accelerate or slow discipline formation?

  3. Failure Modes: Which threat (capture, co-optation, fragmentation, funding drought, elitism) seems most dangerous for Archaeobytology? How would you defend against it?

  4. Case Study Lessons: Should Archaeobytology follow the DH model (slow academic legitimation), Data Science model (fast industry-driven growth), or STS model (interdisciplinary coalition)? Or some hybrid?

  5. Coalitions: Who else should be allied with Archaeobytology that wasn’t mentioned? What organizations or movements should we partner with?

  6. Action Plan: What’s one concrete thing you’ll do in the next month to help build Archaeobytology as a discipline?


Exercise: Draft Your Movement Strategy §

Task: You’re leading the Archaeobytology movement. Design a 5-year strategic plan.

Part 1: Situation Analysis (500 words)

  • Current state (2025): What infrastructure exists?
  • SWOT analysis: Strengths, Weaknesses, Opportunities, Threats
  • Key stakeholders: Who cares about this work?

Part 2: Goals and Metrics (500 words)

For each dimension, set 5-year goals:

  • Knowledge: (journals, conferences, textbooks)
  • Institutions: (programs, centers, departments)
  • Careers: (jobs, certification, placements)
  • Visibility: (media, books, public awareness)
  • Policy: (legislation, testimony, advocacy wins)

Include measurable metrics (e.g., “3 universities with certificates” not just “more programs”)

Part 3: Priority Actions (1000 words)

Choose 10 highest-priority actions for Years 1-5:

  • What should happen first? (Sequence matters)
  • Who leads each action? (students, practitioners, professors, administrators)
  • What resources needed? (funding, staff, space, technology)
  • How to measure success?

Part 4: Risk Mitigation (500 words)

  • What could go wrong?
  • Contingency plans for each failure mode
  • How to stay on track if funding dries up, key people leave, or external crises happen?

Part 5: Call to Action (300 words)

  • If you published this plan publicly, how would you recruit people?
  • What’s the rallying cry?
  • How do you inspire collective action?

Further Reading §

On Discipline Formation §

  • Abbott, Andrew. Chaos of Disciplines. University of Chicago Press, 2001.
  • Klein, Julie Thompson. Interdisciplining Digital Humanities. University of Michigan Press, 2015.
  • Small, Mario Luis. “How to Conduct a Mixed Methods Study.” Annual Review of Sociology 37 (2011): 57-86.

On Movement Building §

  • Ganz, Marshall. “Why David Sometimes Wins: Leadership, Organization, and Strategy in the California Farm Worker Movement.” Oxford, 2009.
  • McAdam, Doug, and Ronnelle Paulsen. “Specifying the Relationship Between Social Ties and Activism.” American Journal of Sociology 99, no. 3 (1993): 640-667.
  • Staggenborg, Suzanne. “The Consequences of Professionalization and Formalization in the Pro-Choice Movement.” American Sociological Review (1988): 585-605.

On Academic Coalition Building §

  • Star, Susan Leigh, and James Griesemer. “Institutional Ecology, ‘Translations’ and Boundary Objects.” Social Studies of Science 19, no. 3 (1989): 387-420.
  • Frickel, Scott, and Neil Gross. “A General Theory of Scientific/Intellectual Movements.” American Sociological Review 70, no. 2 (2005): 204-232.

On Professional Pathways §

  • Nowviskie, Bethany. “On the Origin of ‘Hack’ and ‘Yack.’” In Debates in the Digital Humanities, 2012.
  • Posner, Miriam. “Here and There: Creating DH Community.” In Debates in the Digital Humanities 2016, 2016.

Primary Sources §

  • Archive Team. https://archiveteam.org
  • 4S (Society for Social Studies of Science). https://www.4sonline.org
  • ADHO (Alliance of Digital Humanities Organizations). https://adho.org
  • Society of American Archivists. https://www2.archivists.org

End of Chapter 16 — End of Part IV: Systems & Movements

Next: Part V — Public Scholarship & The Future Chapter 17 — The Public Intellectual in Archaeobytology

Part IV • Disciplinary Movement & The Post-Platform Future

Chapter 17: The Public Intellectual in Archaeobytology

24 min read 5,200 words

Opening: The Scholar in the Arena §

In 2012, Rebecca Solnit wrote an essay called “Men Explain Things to Me” for Guernica magazine. It went viral, spawning the term “mansplaining” and igniting conversations about gender, power, and communication. The essay was accessible, sharp, and personal—nothing like an academic paper.

Solnit is a scholar (cultural historian, essayist). But she’s also a public intellectual—someone who translates complex ideas into public discourse, influencing not just academics but millions of readers, activists, and policymakers.

Archaeobytology needs public intellectuals. Here’s why:

The Problem:

  • Academic papers reach 10-100 scholars
  • Platform shutdowns affect millions of people
  • Policy changes require public pressure
  • Discipline legitimacy requires public visibility

The Gap:

  • Most scholars write for other scholars (jargon-heavy, peer-reviewed journals)
  • Most activists write for activists (insider language, assumed context)
  • The public—voters, users, journalists, politicians—gets neither

The Opportunity: If Archaeobytologists can translate our research into op-eds, podcasts, testimony, and books, we can:

  • Influence policy (right to archive, platform accountability, data portability)
  • Shape culture (make “digital sovereignty” a household concept)
  • Build legitimacy (public intellectuals make disciplines real)
  • Recruit talent (students discover Archaeobytology through public writing)

This chapter teaches you how to become a public intellectual—not instead of being a scholar, but in addition to it. You’ll learn:

  • How to write for different audiences (academic, practitioner, policy, public)
  • How to engage with media (op-eds, podcasts, TV)
  • How to build a platform (blog, newsletter, social media)
  • How to influence policy (testimony, whitepapers, advocacy)
  • How to balance public work with academic expectations (tenure, credibility)

By the end, you’ll have a 5-year strategy for translating your Archaeobytology work into public impact.


Part I: The Five Skills of Public Intellectuals §

Skill 1: Writing for Different Audiences §

Academic writing has its place. But to reach the public, you must write differently.

Audience Matrix

Audience Venue Length Tone Evidence Goal
Academic Peer-reviewed journals 8,000-12,000 words Formal, cautious Exhaustive citations Advance knowledge
Practitioners Trade publications 2,000-3,000 words Professional, actionable Case studies Improve practice
Policymakers White papers, briefs 1,000-1,500 words Clear, evidence-based Key statistics, recommendations Inform decisions
General Public Op-eds, magazines 800-1,200 words Accessible, urgent Stories + data Shape discourse
Social Media Twitter threads, LinkedIn 200-500 words Conversational, shareable Hooks, visuals Start conversations

Translation Exercise: One Idea, Five Audiences

Academic Version (for Journal of Archaeobytology):

“The phenomenon of platform-mediated digital mortality—wherein corporate entities terminate hosting infrastructure, resulting in the permanent deletion of user-generated content—represents a novel form of cultural erasure distinct from traditional archival loss. Unlike material artifacts, which decay gradually and leave archaeological traces, digital artifacts experience catastrophic failure: the transition from accessibility to permanent inaccessibility occurs instantaneously upon server decommission.”

Practitioner Version (for Library Journal):

“When platforms shut down, libraries face a new challenge: digital content doesn’t decay slowly like books—it vanishes overnight. This ‘platform death’ requires proactive archiving strategies. Librarians must scrape endangered platforms before shutdown, not wait for donation of already-lost materials.”

Policy Version (for Congressional brief):

“Platform shutdowns have deleted billions of cultural artifacts, including historical documentation of social movements, journalism, and community organizing. Recommendation: Mandate 90-day notice for platform shutdowns + require user data export in open formats. Cost to industry: minimal. Benefit to cultural preservation: substantial.”

Public Version (for New York Times op-ed):

“When GeoCities died in 2009, 30 million websites vanished overnight. Your teenage homepage. Your friend’s memorial site. An entire era of internet culture—murdered by Yahoo with three weeks’ notice. This wasn’t obsolescence. It was execution. And it keeps happening.”

Social Media Version (Twitter thread):

“🧵 Why do our digital memories keep disappearing?

1/ When GeoCities shut down in 2009, 30M websites died in one day

2/ Not because of technical failure—but because Yahoo decided they weren’t profitable

3/ This is ‘platform murder’—and it’s accelerating

[Thread continues with solutions, call to action]”

Writing Rules by Audience

For Academics:

  • Engage with literature (cite extensively)
  • Be precise, even if verbose
  • Caveat everything (acknowledge limitations)
  • Original data/methods required

For Practitioners:

  • Lead with the problem (they need solutions)
  • Provide actionable steps (checklists, workflows)
  • Use real examples (case studies they recognize)
  • Skip theory unless it improves practice

For Policymakers:

  • Executive summary first (they won’t read past page 1 if not hooked)
  • Use numbers (cost-benefit, impact metrics)
  • Clear recommendations (numbered list, specific actions)
  • Bipartisan framing (“preserving culture” not “regulating tech”)

For Public:

  • Start with story (not theory)
  • Use everyday language (no jargon)
  • Make it urgent (why should reader care now?)
  • End with action (what can they do?)

For Social Media:

  • Hook in first sentence (make them click “read more”)
  • One idea per post (don’t pack in everything)
  • Visual anchors (images, charts, GIFs)
  • Invite engagement (ask questions, request replies)

Skill 2: Media Engagement §

Journalists are megaphones. If you can work with them, your ideas reach millions.

Types of Media Engagement

1. Reactive Commentary (Breaking News)

  • Journalist needs expert quote on breaking story
  • Example: “Twitter announces shutdown—expert comments?”
  • Timeline: Hours (respond same day or lose opportunity)
  • Format: 2-3 quotable sentences
  • How to prepare: Monitor news, respond quickly, have pre-written points

2. Feature Interviews

  • Journalist writing longer piece, interviews you as expert
  • Example: Wired feature on digital preservation
  • Timeline: Days to weeks
  • Format: 30-60 min interview, journalist selects quotes
  • How to prepare: Provide stories, data, concrete examples; offer to connect them with other sources

3. Op-Ed Pitching

  • You write opinion piece, pitch to publication
  • Example: “Why Platform Shutdowns Are Cultural Violence” for The Atlantic
  • Timeline: Write first, pitch immediately after news hook
  • Format: 800-1,200 words, strong argument
  • How to prepare: Study publication’s style, find news hook, pitch editor with strong lede

4. Podcast Appearances

  • Interview on podcast (30-90 min conversation)
  • Example: Reply All, On The Media, The Ezra Klein Show
  • Timeline: Book weeks in advance, record for hour, edited to 30-45 min
  • Format: Conversational, tell stories, explain concepts
  • How to prepare: Listen to past episodes, prepare 3-5 key points, practice telling stories

5. TV/Video

  • News segments, documentaries
  • Example: CNN interview on platform accountability
  • Timeline: Often last-minute (breaking news) or months (documentaries)
  • Format: 3-5 min segments (TV), longer (documentaries)
  • How to prepare: Master the sound bite (7-10 second quotable points), dress professionally, speak in complete sentences (no “um,” “like”)

Building Media Relationships

Create a Media Kit:

  • Bio: 2-3 sentences (who you are, your expertise)
  • Headshot: Professional photo (300dpi)
  • Expertise list: Topics you can speak on (bullet points)
  • Past media: Links to previous interviews, op-eds
  • Contact: Email, phone (make it easy for journalists to reach you)

Be Responsive:

  • Journalists have tight deadlines (hours, not days)
  • If you can’t respond immediately, refer them to colleague who can
  • Reply even if you can’t help (“I don’t know this, but Dr. X does—here’s their email”)

Offer More Than Asked:

  • Journalist asks for quote → offer to send data, visuals, or other expert contacts
  • Journalist asks about X → mention related angle Y they might not have considered
  • Build reputation as helpful source (they’ll come back)

Pitch Proactively:

  • Don’t wait to be asked—pitch ideas to journalists
  • Example: “Hi [Journalist], I follow your tech coverage. I have data on platform shutdowns that might interest you for a feature. Would you like to see the findings?”
  • Target journalists who cover your beat (study their past work)

Case Study: Safiya Noble’s Media Strategy

Background: Safiya Noble is a scholar (USC professor) who wrote Algorithms of Oppression (2018) about racist search engine results.

Media Trajectory:

  1. Academic foundation: Published peer-reviewed research
  2. Public book: Translated research into accessible book (Algorithms of Oppression)
  3. Op-eds: Wrote for The Guardian, The Washington Post connecting research to breaking news
  4. Congressional testimony: Invited to testify on algorithmic bias (2019)
  5. Documentary appearances: Featured in Coded Bias film
  6. Mainstream visibility: CNN, NPR, New York Times interview her as go-to expert

Timeline: 10+ years from first research to mainstream recognition

Strategy:

  • Built academic credibility first (peer review, tenure)
  • Wrote accessible book (not just articles)
  • Seized news hooks (Google autocomplete scandals)
  • Cultivated journalist relationships (responded quickly, provided data)

Result: Policy impact (companies changed algorithms), cultural shift (“algorithmic bias” entered public discourse), discipline building (helped establish field).

Skill 3: Public Speaking §

Speaking is different from writing. You must hold attention, adapt in real-time, and connect emotionally.

Speaking Venues for Archaeobytologists

1. Academic Conferences

  • Format: 20-min paper + Q&A
  • Audience: Scholars (knowledgeable, critical)
  • Goal: Advance knowledge, get feedback, network
  • Style: Formal, data-driven, lit review

2. Industry Keynotes

  • Format: 30-45 min talk (TEDx-style)
  • Audience: Practitioners (want actionable insights)
  • Goal: Inspire, provide frameworks, establish authority
  • Style: Story-driven, visual slides, “takeaways”

3. TED/TEDx Talks

  • Format: 18 min max, memorized, no notes
  • Audience: General public (curious, diverse backgrounds)
  • Goal: Spread one big idea, go viral
  • Style: Narrative arc, emotional connection, no jargon

4. Policy Hearings/Testimony

  • Format: 5 min prepared statement + Q&A
  • Audience: Legislators, staffers (busy, need summaries)
  • Goal: Inform policy, establish credibility
  • Style: Evidence-based, clear recommendations, respectful

5. Public Lectures

  • Format: 45-60 min + Q&A
  • Audience: General public (educated, interested)
  • Goal: Educate, provoke thought, recruit allies
  • Style: Accessible but substantive, Q&A is crucial

6. Podcasts (Interviewed)

  • Format: 45-90 min conversation
  • Audience: Niche (listeners of that podcast)
  • Goal: Deep dive, build following, humanize research
  • Style: Conversational, tell stories, be yourself

The Rule of Three

People remember three things from a talk. No more. Design around this.

Bad Talk Structure: “I’ll discuss 7 dimensions of platform death, 12 preservation methods, and 15 policy recommendations.”

Good Talk Structure: “Three reasons platforms murder culture:

  1. Profit (you’re not profitable anymore)
  2. Control (you’re not controllable anymore)
  3. Liability (you’re a legal risk now)

And three things we can do:

  1. Archive (save it before it dies)
  2. Build alternatives (so we’re not hostage)
  3. Legislate (make murder harder)”

Audience remembers: Profit/Control/Liability + Archive/Build/Legislate

Speaking Best Practices

1. Start with Story, Not Theory

  • Bad: “Today I’ll discuss the theoretical framework of platform mortality.”
  • Good: “In 2009, my childhood GeoCities page vanished overnight. Yahoo deleted 30 million websites with three weeks’ notice. This is why…”

2. Show, Don’t Tell

  • Don’t describe a GeoCities page—show a screenshot
  • Don’t explain “platform murder”—play a video of deleted content
  • Visuals > words

3. Practice Out Loud

  • Reading your talk silently ≠ speaking it
  • Record yourself, listen back, identify awkward phrasing
  • Time yourself (audiences hate talks that run over)

4. Prepare for Q&A

  • Anticipate hostile questions (“Isn’t this just nostalgia for old tech?”)
  • Have 2-3 prepared answers to predictable questions
  • It’s okay to say “I don’t know” (better than bullshitting)

5. Make It Interactive

  • Ask audience questions (“How many of you have lost digital content to platform shutdowns?”)
  • Invite participation (“Turn to the person next to you and discuss…”)
  • Q&A should be dialogue, not interrogation

Skill 4: Platform Building §

Public intellectuals need platforms—ways to reach audiences directly, not mediated by institutions or publications.

Platform Options

Platform Time Investment Reach Potential Control Longevity
Personal Blog High (3-5 hrs/post) Low→High (SEO growth) Total Decades (if you own domain)
Newsletter (Substack, Ghost) Medium (1-2 hrs/issue) Medium (subscriber growth) High Years (portable)
Twitter/Mastodon Medium (30 min/day) High (viral potential) Low (platform controls) Uncertain (platform risk)
YouTube Very High (full production) Very High (algorithm boost) Low (platform controls) Years (but platform-dependent)
Podcast High (recording + editing) Medium Medium Years
LinkedIn Low (15 min/post) Medium (professional network) Low Years (stable platform)
TikTok Medium (short videos) Very High (algorithm favors new creators) Low Uncertain

Cory Doctorow’s “Pluralistic” Model (Gold Standard)

What He Does:

  • Daily blog post (1,000-3,000 words)
  • Cross-posts to:
  • His blog (pluralistic.net—he owns domain)
  • Twitter (threaded)
  • Mastodon (federated, sovereignty!)
  • Tumblr (visual platform)
  • Weekly newsletter compiling week’s posts

Why It Works:

  • Consistency: Daily output builds audience
  • Sovereignty: Owns domain (pluralistic.net), so if platforms die, content persists
  • Reach: Cross-posting gets content to multiple audiences
  • Redundancy: If one platform bans him, others remain

Result: 100,000+ readers, major influence on tech policy (cited in EU Digital Markets Act), financially sustainable (book deals, speaking fees), didn’t require institutional backing.

Lessons:

  1. Own your domain (yourname.com → foundation of platform)
  2. Consistency > volume (daily 1,000 words > weekly 7,000 words)
  3. Cross-post strategically (reach people where they are, but keep original on your site)
  4. Build email list (social media can ban you; email is yours)

Building Your Platform: 5-Year Plan

Year 1: Establish Foundation

  • Register domain (yourname.com)
  • Set up blog (WordPress, Ghost, or static site)
  • Commit to frequency (weekly is realistic for most academics)
  • Topics: Your research, accessible explanations, responses to news

Year 2: Grow Audience (0 → 1,000 readers)

  • SEO optimization (write about topics people search for)
  • Guest post on established sites (borrow audiences)
  • Engage on social media (share your posts, comment on others’)
  • Email list: Offer newsletter signup (target: 100 subscribers by year end)

Year 3: Diversify Platforms (1,000 → 5,000 readers)

  • Add second platform (newsletter, podcast, or video)
  • Cross-promote (blog readers → newsletter subscribers)
  • Collaborate (interview other experts, be interviewed)
  • Speaking: Accept invitations, build reputation

Year 4: Establish Authority (5,000 → 10,000+ readers)

  • Publish book (grows credibility, reaches new audiences)
  • Media appearances (journalists find you via blog, invite you for quotes)
  • Keynote speeches (paid speaking opportunities)
  • Consider monetization (Patreon, paid newsletter tier, consulting)

Year 5: Sustainable Impact (10,000+ readers)

  • Platform is established (people know your name)
  • Media regularly quotes you (go-to expert)
  • Policy influence (testimony, advisory roles)
  • Can support yourself partly/fully from platform (if desired)

Reality Check:

  • Not everyone reaches 10,000 readers (that’s okay!)
  • Even 500 engaged readers = influence (if they’re decision-makers, journalists, other scholars)
  • Platform building is long game (think years, not months)

Skill 5: Policy Influence §

Archaeobytologists should shape laws, not just study what laws allow.

Mechanisms of Policy Influence

1. White Papers

  • What: Research-based policy recommendations (10-30 pages)
  • When: Before legislation is drafted (shape the conversation)
  • Example: “A Framework for Right to Archive Legislation”
  • Audience: Policymakers, staffers, advocacy organizations

2. Congressional/Parliamentary Testimony

  • What: Invited to speak at hearing (5 min prepared statement + Q&A)
  • When: When lawmakers are considering relevant legislation
  • Example: Testifying on platform accountability bill
  • Impact: Your testimony becomes part of legislative record

3. Op-Eds in Policy Context

  • What: Opinion piece in major paper, timed to legislative debate
  • When: During policy windows (bill being considered, scandal breaking)
  • Example: “Why Congress Must Mandate Data Portability” in Washington Post
  • Impact: Lawmakers read these; staffers send them to bosses

4. Coalition Letters

  • What: Open letter signed by experts, organizations
  • When: Supporting or opposing specific legislation
  • Example: “100 Scholars Call for Right to Archive Law”
  • Impact: Shows consensus, makes lawmakers pay attention

5. Informal Briefings

  • What: Meeting with staffers to explain complex issues
  • When: Ongoing (build relationships, offer expertise)
  • Example: “Lunch briefing on digital preservation challenges”
  • Impact: Staff learn from you, remember you when drafting bills

6. Model Legislation

  • What: Draft actual legal language for lawmakers to introduce
  • When: When you have clear policy prescription
  • Example: “Digital Right of First Refusal Act” (model bill)
  • Impact: Makes it easy for lawmakers (they can introduce your bill verbatim)

Building Policy Influence: The Pathway

Phase 1: Establish Legitimacy (Years 1-2)

  • Publish research (peer-reviewed, credible)
  • Build public profile (op-eds, speaking)
  • Join relevant organizations (EFF, ALA, advocacy groups)

Phase 2: Get on the Radar (Years 2-4)

  • Write white papers (circulate to policy orgs)
  • Testify at state/local hearings (build testimony experience)
  • Op-eds when relevant bills are debated
  • Meet staffers (offer expertise, don’t demand anything)

Phase 3: Direct Influence (Years 4-10)

  • Congressional testimony (federal level)
  • Draft model legislation (with advocacy partners)
  • Join advisory boards (FCC, FTC, Library of Congress)
  • International work (WIPO, UNESCO, EU)

Case Study: Brewster Kahle’s Policy Strategy

Background: Founder of Internet Archive (1996)

Policy Trajectory:

  1. Built credibility: Internet Archive became indispensable resource (billions of archived pages)
  2. Legal advocacy: Fought for library lending rights (controlled digital lending)
  3. Coalition building: Partnered with libraries, scholars, advocacy groups
  4. Public visibility: TED talks, interviews, positioned as “librarian of the internet”
  5. Direct testimony: Testified before Congress on copyright, preservation, access
  6. Model proposals: Advanced proposals like “Digital Public Library of America”

Impact:

  • Internet Archive’s practices influenced copyright policy debates
  • Positioned digital preservation as public interest (not just technical hobby)
  • Made “universal access to knowledge” a mainstream policy goal

Lessons:

  • Build something valuable first (gives you standing)
  • Frame issues broadly (public interest, not narrow technical concerns)
  • Partner with established institutions (libraries, universities)
  • Play long game (decades of advocacy, not one-off campaigns)

Part II: Balancing Public and Academic Work §

The Tenure Trap §

The Problem:

  • Tenure committees value peer-reviewed articles (op-eds don’t count)
  • Public work takes time away from research (opportunity cost)
  • Some academics view public intellectuals as “popularizers” (not serious scholars)

The Risk:

  • Pre-tenure faculty write op-eds → denied tenure (“didn’t publish enough”)
  • Public visibility threatens academic credibility (“too political,” “not rigorous”)

The Reality: This is changing. Slowly. Some fields now value “public scholarship” (especially in humanities). But risk remains.

Strategies for Pre-Tenure Faculty

1. Prioritize Peer Review First

  • Get articles published in top journals (establish academic credentials)
  • Public work is supplement, not substitute
  • Rule of thumb: 1 public piece for every 2 academic articles

2. Frame Public Work as Impact

  • In tenure file, argue that op-eds/testimony demonstrate research impact
  • Show citations (your op-ed cited by policymakers, journalists)
  • Quantify reach (10,000 readers vs. 100 for academic article)

3. Choose Safe Venues

  • Chronicle of Higher Education, Inside Higher Ed (academic-adjacent publications)
  • University press trade books (peer-reviewed but accessible)
  • Public scholarship journals (Public Historian, Engaging Science, Technology, and Society)

4. Get Support from Senior Colleagues

  • Find mentors who value public work
  • Ask them to write letters emphasizing importance of engagement
  • Form alliances with like-minded faculty

5. Know Your Institution

  • R1 universities: Prioritize traditional research (play it safe pre-tenure)
  • Teaching-focused colleges: May value public engagement more
  • Ask during job interview: “Does this department value public scholarship?”

Post-Tenure Freedom

Once tenured, you have more freedom:

  • Public work can’t hurt you (job security)
  • You’ve proven academic credibility (can now experiment)
  • Platform from tenure gives you authority (journalists want credentialed experts)

Many scholars become public intellectuals after tenure:

  • Spend years 1-6 publishing in journals
  • Get tenure at year 6-7
  • Years 8+ shift toward op-eds, books, testimony

This is a viable path. Play the academic game first, then use that credibility for public impact.


Part III: Five-Year Public Intellectual Strategy §

Let’s design your personal roadmap.

Your Strategy Canvas §

Fill this out to create your plan:

1. Personal Brand (Who Are You?)

Your Niche:

  • What’s your specific expertise within Archaeobytology?
  • Example: “Digital preservation of social movements” or “Platform governance and user rights”

Your Elevator Pitch:

  • 2-3 sentences: Who you are, what you study, why it matters
  • Example: “I’m an Archaeobytologist studying how platforms murder culture. When GeoCities died, 30 million websites vanished. I preserve endangered platforms and build alternatives that can’t be killed.”

Your Unique Angle:

  • What do you bring that others don’t?
  • Example: “I’m the only person studying early trans YouTube comprehensively” or “I combine legal expertise with technical preservation skills”

2. Platform Strategy (Where Will You Publish?)

Primary Platform (Where you’ll invest most time):

  • Blog? Newsletter? YouTube? Podcast?
  • Choose one, own it, make it yours

Secondary Platform (For distribution):

  • Social media (Twitter, Mastodon, LinkedIn)
  • Cross-post links to drive traffic to primary

Growth Goal:

  • Year 1: 0 → 100 readers/subscribers
  • Year 3: 100 → 1,000
  • Year 5: 1,000 → 5,000+

3. Writing Strategy (Academic + Public)

Academic Track (For tenure/credibility):

  • Target: 2-3 peer-reviewed articles per year
  • Journals: Journal of Archaeobytology, Digital Humanities Quarterly, Social Studies of Science

Public Track (For impact/visibility):

  • Target: 6-12 op-eds/blog posts per year (monthly or biweekly)
  • Venues: Chronicle of Higher Education, Wired, The Atlantic, your blog

Book Plan:

  • Years 1-3: Article publications (build CV)
  • Years 4-5: Write book (synthesize research for broader audience)
  • Year 6+: Book publication → media tour, speaking invitations

4. Speaking Strategy (Local → National → High-Profile)

Year 1-2: Local

  • University talks (your dept, other depts on campus)
  • Local libraries, community groups
  • Goal: Practice speaking, refine message

Year 3-4: National

  • Academic conferences (SAA, ADHO, 4S)
  • Industry conferences (tech conferences, library associations)
  • Podcasts (pitch yourself as guest)
  • Goal: Build network, get known in field

Year 5+: High-Profile

  • TEDx talks (apply for speaking slots)
  • Congressional testimony (through advocacy organizations)
  • Major media (NPR, CNN when your issue is in news)
  • Goal: Reach mass audiences, shape policy

5. Media Strategy (Reactive + Proactive)

Media Kit (Create in Year 1):

  • Bio (200 words)
  • Headshot (professional photo)
  • Expertise list (“I can speak on: platform shutdowns, digital preservation, user rights”)
  • Past media (links to any interviews, op-eds)
  • Contact (email, phone)

Proactive Pitching (Ongoing):

  • Identify 5-10 journalists who cover your beat
  • Follow them on social media, read their work
  • Pitch ideas when you have data or timely hook
  • Example: “Hi [journalist], I follow your tech coverage. I just released data on platform shutdowns that shows X trend. Would you be interested in covering this?”

Responsive Engagement (When news breaks):

  • Monitor news for stories related to your work
  • Reply quickly when journalists request expert comment (within hours)
  • Offer more than asked (data, visuals, other expert contacts)

6. Policy Strategy (Legitimacy → Testimony → Legislation)

Year 1-2: Build Legitimacy

  • Publish research (establish expertise)
  • Join organizations (EFF, ALA, advocacy groups)
  • Write white papers (share with policy orgs)

Year 3-4: Get Invited

  • Testify at local/state hearings
  • Write op-eds when bills are debated
  • Informal briefings with staffers

Year 5+: Direct Influence

  • Congressional testimony
  • Draft model legislation (with partners)
  • Advisory roles (FCC, FTC, etc.)

7. Impact Metrics (How Will You Know You’re Succeeding?)

Quantitative:

  • Website/blog traffic (pageviews, unique visitors)
  • Email subscribers
  • Social media followers
  • Speaking invitations
  • Media mentions (times you’re quoted)
  • Policy citations (your work cited in testimony, reports)

Qualitative:

  • Recognition (people in your field know your name)
  • Influence (your ideas show up in others’ work, policy debates)
  • Community (you’ve built network of allies, collaborators)
  • Cultural impact (concepts you coined enter public discourse)

8. Three Pillars Check (Are You Sovereign?)

  • Declaration: Do you own your platform? (Your domain, not Medium)
  • Connection: Can you reach your audience directly? (Email list, not just Twitter)
  • Ground: Do you control your content? (Local backups, exportable formats)

If you’re building public platform on someone else’s land (Medium, Substack), you’re vulnerable. Aim for sovereignty.

9. Risk Analysis (What Could Go Wrong?)

Burnout:

  • Risk: Public work + academic work = too much
  • Mitigation: Set boundaries (e.g., “I write one op-ed per month, no more”)

Backlash:

  • Risk: Public visibility invites criticism, harassment
  • Mitigation: Don’t read comments, have support network, know when to step back

Co-optation:

  • Risk: Media oversimplifies your work, misrepresents your views
  • Mitigation: Insist on reviewing quotes, clarify when misquoted, own your platform (blog) to set record straight

Institutional Pushback:

  • Risk: Tenure committee doesn’t value public work
  • Mitigation: Prioritize peer review pre-tenure, frame public work as impact

Time Sink:

  • Risk: Public work takes time from research, teaching, life
  • Mitigation: Be strategic (one great op-ed > ten mediocre tweets), batch work (write multiple pieces at once)

Conclusion: From Scholar to Public Figure §

Public intellectual work is not a betrayal of scholarship—it’s an extension. You’re taking the knowledge you create and making it matter beyond the academy.

The Goal Is Not Fame: It’s influence. You want your ideas to shape:

  • Policy (laws that protect digital culture)
  • Culture (concepts that change how people think)
  • Practice (methods that others adopt)
  • Discipline (Archaeobytology becomes real)

The Path Is Long: 5-10 years to go from “nobody knows me” to “go-to expert.” But every op-ed, every talk, every testimony moves you forward.

The Work Is Necessary: Archaeobytology can’t become legitimate if it stays in academic journals. We need public intellectuals who can:

  • Explain platform death to New York Times readers
  • Testify before Congress on digital rights
  • Write books that students discover and think “I want to study this”
  • Build platforms that demonstrate digital sovereignty in practice

In the next chapter—the final chapter—we’ll bring it all together: forging the Third Way, the vision for a post-platform future, and the Archaeobytologist’s Manifesto.

But first, consider: What’s your public intellectual strategy? If you spent the next five years building a platform, engaging media, and influencing policy, where would you be? And what would Archaeobytology as a field gain?

The discipline needs scholars. But it also needs public intellectuals.

Will you be one?


Discussion Questions §

  1. Personal Assessment: Do you see yourself as a public intellectual, or purely an academic? Why? What appeals to or scares you about public work?

  2. Tradeoffs: How do you balance scholarly rigor with public accessibility? Where’s the line between “simplified explanation” and “oversimplification”?

  3. Platform Choice: Which platform would you choose for public intellectual work? (Blog, newsletter, podcast, social media, video?) What drives your choice?

  4. Risk Tolerance: How much professional risk are you willing to take? Would you write controversial op-eds pre-tenure? Or wait until tenured?

  5. Role Models: Who are public intellectuals you admire (in any field)? What do they do well? What would you do differently?

  6. Impact Metrics: How would you measure success? Follower counts? Policy citations? Just “more people understand Archaeobytology”?


Exercise: Draft Your 5-Year Public Intellectual Strategy §

Task: Create a concrete plan for becoming a public intellectual in Archaeobytology.

Part 1: Brand and Positioning (500 words) §

  • Niche: What’s your specific expertise?
  • Elevator pitch: Who are you, what do you do, why does it matter? (3 sentences)
  • Unique angle: What do you bring that others don’t?
  • Target audiences: Who do you want to reach? (Academics, practitioners, policymakers, general public?)

Part 2: Platform Strategy (500 words) §

  • Primary platform: Where will you publish? (Blog, newsletter, YouTube, podcast?)
  • Domain: Will you own yourname.com or use a hosted platform?
  • Posting frequency: How often can you realistically create content?
  • Content types: What will you write/create about?
  • Growth plan: How will you build audience (0 → 100 → 1,000 → 5,000)?

Part 3: Media Engagement (500 words) §

  • Media kit: Draft your 200-word bio, expertise list, contact info
  • Target journalists: List 5-10 journalists who cover your beat
  • Pitch strategy: How will you get their attention?
  • Op-ed ideas: Brainstorm 3 op-ed topics with news hooks

Part 4: Speaking and Policy (500 words) §

  • Speaking venues: Where will you speak (Years 1-2 vs. Years 3-5)?
  • Policy pathway: How will you influence policy? White papers? Testimony? Advocacy partnerships?
  • Key messages: What are the 3 core ideas you want policymakers to understand?

Part 5: Risk Mitigation (300 words) §

  • Burnout prevention: How will you avoid overcommitting?
  • Tenure strategy: If pre-tenure, how will you balance public/academic work?
  • Backlash plan: How will you handle criticism, harassment?
  • Boundaries: What will you say no to?

Part 6: Timeline and Milestones (300 words) §

Create a 5-year timeline with concrete milestones:

  • Year 1: Launch blog, publish 3 op-eds, give 2 local talks
  • Year 2: Grow to 500 subscribers, appear on 2 podcasts, testify at local hearing
  • Year 3: Write book proposal, publish 6 op-eds, speak at national conference
  • Year 4: Book published, media tour, congressional testimony
  • Year 5: Established expert, 5,000+ readers, advisory role

Part 7: Reflection (200 words) §

  • Excitement: What excites you most about this plan?
  • Fear: What scares you?
  • Feasibility: Is this realistic given your life circumstances?
  • Commitment: Will you actually do this? Why or why not?

Further Reading §

On Public Scholarship §

  • Burawoy, Michael. “For Public Sociology.” American Sociological Review 70, no. 1 (2005): 4-28.
  • Manifesto for scholars engaging beyond academy

  • Posner, Miriam. “What’s Next: The Radical, Unrealized Potential of Digital Humanities.” In Debates in the Digital Humanities 2016, edited by Matthew Gold and Lauren Klein, 32-41. University of Minnesota Press, 2016.

  • DH scholar on making scholarship matter

On Public Intellectuals §

  • Jacoby, Russell. The Last Intellectuals: American Culture in the Age of Academe. Basic Books, 1987.
  • Classic (pessimistic) account of public intellectuals’ decline

  • Small, Helen. The Value of the Humanities. Oxford University Press, 2013.

  • Defending humanities in public sphere

On Writing for Public §

  • Sword, Helen. Stylish Academic Writing. Harvard University Press, 2012.
  • How to write accessibly without dumbing down

  • Pinker, Steven. The Sense of Style. Viking, 2014.

  • Cognitive science of clear writing

Case Studies §

  • Noble, Safiya Umoja. Algorithms of Oppression. NYU Press, 2018.
  • Example of scholarship → public impact

  • Doctorow, Cory. “Pluralistic.” https://pluralistic.net/

  • Daily blog, model for public intellectual platform

On Media Engagement §

  • Nisbet, Matthew, and Dietram Scheufele. “What’s Next for Science Communication? Promising Directions and Lingering Distractions.” American Journal of Botany 96, no. 10 (2009): 1767-1778.
  • How scientists engage media (applicable to all scholars)

On Policy Influence §

  • Pielke, Roger. The Honest Broker. Cambridge University Press, 2007.
  • How scientists influence policy (without becoming advocates)

End of Chapter 17

Next: Chapter 18 — Forging the Third Way: Vision for a Post-Platform Future (The final chapter! The manifesto!)

Part IV • Disciplinary Movement & The Post-Platform Future

Chapter 18: Forging the Third Way

Vision for a Post-Platform Future

22 min read 4,695 words

Opening: The Crossroads §

We stand at a crossroads in the history of digital culture.

Path 1: Platform Feudalism

  • Continued consolidation under Big Tech monopolies
  • Users as perpetual tenants, renting digital existence
  • Culture murdered whenever it’s unprofitable
  • Surveillance capitalism extracting behavioral data as raw material
  • Every generation loses its digital history to corporate whims

Path 2: Regulatory Containment

  • Governments regulate platforms (antitrust, interoperability mandates, data protection)
  • Platforms become quasi-utilities, like phone companies
  • Improvement over feudalism, but still centralized
  • Corporate landlords remain, just with more oversight
  • Users gain some protections but not sovereignty

Path 3: Digital Sovereignty (The Third Way)

  • Users own their identities, connections, and ground
  • Distributed infrastructure: federated, P2P, cooperative
  • Culture persists independent of corporate survival
  • Economic models that don’t require surveillance or extraction
  • Preservation built into system design, not emergency afterthought

This book has been building toward Path 3. Every chapter—from the Archaeobyte Taxonomy to the Three Pillars, from Triage to Institution Building, from Movement Strategy to Public Intellectual practice—has prepared you to forge the Third Way.

This final chapter asks: What does the Third Way actually look like? Not as abstract ideal, but as concrete system design. What would we build if we started over, knowing everything we know about how platforms murder culture?

This is our manifesto. Our blueprint. Our declaration that another internet is possible.


Part I: Principles of the Third Way §

Before designing systems, we must articulate core principles—the non-negotiables that distinguish the Third Way from both feudalism and containment.

Principle 1: User Sovereignty Is Non-Negotiable §

Declaration, Connection, Ground must be user-owned, not platform-granted.

This means:

  • Identities are portable: [email protected], not platform.com/you
  • Data is exportable: Full archives, usable formats, no lock-in
  • Infrastructure is exit-able: Can migrate between providers without losing connections

What this rules out:

  • Platforms that own your username
  • Social graphs you can’t export
  • Proprietary formats that trap your data

What this enables:

  • Federation (Mastodon, Matrix, email model)
  • Self-hosting (for those with technical capacity)
  • Portable hosting (Ghost, WordPress—custom domain, full export)

Principle 2: Preservation Is a Design Constraint, Not an Afterthought §

Systems must be built to outlast their creators.

This means:

  • Open standards: Protocols anyone can implement (not proprietary APIs)
  • Documented architectures: Future archaeologists can understand how it worked
  • Redundant storage: LOCKSS principle (Lots of Copies Keep Stuff Safe)
  • Graceful degradation: If advanced features fail, basic content remains accessible

What this rules out:

  • Closed-source platforms with no documentation
  • Centralized servers as single point of failure
  • Formats that require vendor software to read

What this enables:

  • Internet Archive can crawl and preserve
  • Community can fork if maintainers abandon
  • Content survives platform death

Principle 3: Surveillance Capitalism Is Incompatible with Sovereignty §

You cannot be sovereign if platforms monetize your behavior through surveillance.

This means:

  • No behavioral tracking for ads: No surveillance infrastructure
  • Transparent business models: Users know how platform makes money
  • Data minimization: Collect only what’s needed for service to function

What this rules out:

  • Facebook/Google ad model (surveillance-funded)
  • “Free” services that sell user data
  • Algorithmic manipulation for engagement (rage-farming)

What this enables:

  • Subscriptions (Ghost, Fastmail)
  • Freemium (Proton, Signal)
  • Cooperatives (user-owned platforms)
  • Public funding (Wikipedia, NPR model)

Principle 4: Interoperability Over Monopoly §

Network effects must not create lock-in.

This means:

  • Open protocols: ActivityPub, Matrix, RSS—anyone can implement
  • Account portability: Can switch providers, keep followers
  • Cross-platform communication: Email model (Gmail users can email Outlook users)

What this rules out:

  • Walled gardens (Instagram can’t message TikTok)
  • Platform-specific features that prevent migration
  • Proprietary networks with no bridges

What this enables:

  • Competition (switching costs are low)
  • Innovation (anyone can build better client)
  • Exit rights (leave bad platform without losing community)

Principle 5: Governance Must Be Democratic, Not Corporate §

Users must have voice in how platforms are run.

This means:

  • Cooperative ownership: Users vote on major decisions
  • Transparent governance: Public board meetings, documented policies
  • Community moderation: Federated model where instance admins set rules

What this rules out:

  • Benevolent dictators (even well-meaning founders eventually sell or die)
  • Venture capital (investors demand growth and exit, not sustainability)
  • Opaque ToS changes (platforms changing rules without user input)

What this enables:

  • Platform cooperatives (Stocksy, Resonate)
  • Federated governance (Mastodon instances)
  • Non-profit stewardship (Wikimedia, Internet Archive)

Principle 6: The Commons Must Be Protected from Enclosure §

Shared cultural resources cannot be privatized.

This means:

  • Public domain by default: Content should eventually enter commons
  • Anti-enclosure licensing: Copyleft (GPL, CC-BY-SA) prevents proprietary capture
  • Archival rights: Society has right to preserve culture, even if corporate copyright opposes

What this rules out:

  • Perpetual copyright (Disney extending terms forever)
  • DRM that prevents preservation
  • Platforms claiming ownership of user-generated content

What this enables:

  • Remix culture (legal to build on others’ work)
  • Long-term preservation (archives can save copyrighted material)
  • Cultural continuity (each generation accesses previous generations’ work)

Part II: System Architecture of the Third Way §

With principles established, how do we build the Third Way? What does the technical architecture look like?

Layer 1: Identity (Declaration) §

Problem: Centralized platforms own your identity. If banned, “you” cease to exist.

Third Way Solution: Federated Identity

Model: Email + Domain Names

  • Your identity: [email protected]
  • Domain is yours (registered, portable)
  • Email provider can change (Gmail → Fastmail → self-hosted), identity stays same

Applied to Social Media:

  • Mastodon: @[email protected]
  • You run instance, or use hosting service (but can migrate)
  • Portable across ActivityPub-compatible platforms

Applied to Authentication:

  • OpenID Connect: yourdomain.com as identity
  • Log in to services with your domain (not “Sign in with Google”)
  • You control authentication (can revoke access)

Key Technologies:

  • DNS (for domain-based identity)
  • ActivityPub (for federated social)
  • DID (Decentralized Identifiers, for blockchain-based identity—though controversial)

Trade-offs:

  • Requires owning domain (~$15/year—barrier for some)
  • Technical complexity higher than creating Facebook account
  • But: True sovereignty requires some cost/effort

Layer 2: Communication (Connection) §

Problem: Platforms mediate all communication, can shadowban, algorithmically filter, or shut down.

Third Way Solution: End-to-End Encrypted, Federated Communication

Model: Email (for public/async) + Signal (for private/sync)

For Public Communication (Posts, Blogs):

  • RSS/Atom: Anyone can subscribe to anyone (no algorithmic feed)
  • ActivityPub: Federated timeline (like email—Gmail users see Outlook users’ posts)
  • Webmentions: Decentralized replies (your blog can reply to mine, no centralized comment system)

For Private Communication (Messaging):

  • Matrix: Federated, E2E encrypted chat (like Signal + email model)
  • Signal Protocol: Gold standard E2E encryption
  • No metadata surveillance: Platforms can’t read content or build social graphs

For Discovery:

  • Search engines: Decentralized (YaCy) or privacy-respecting (DuckDuckGo, Kagi)
  • Social bookmarking: User-curated (not algorithmic)
  • RSS readers: User chooses what to follow (not platform-recommended)

Key Technologies:

  • ActivityPub, Matrix (federation)
  • Signal Protocol (E2E encryption)
  • RSS/Atom (syndication)

Trade-offs:

  • Discovery harder (no algorithmic recommendation of “people you might know”)
  • Requires active curation (following people deliberately, not passively scrolling feed)
  • But: No manipulation, no surveillance

Layer 3: Storage (Ground) §

Problem: Platforms store your data on their servers. If they shut down or ban you, data vanishes.

Third Way Solution: Distributed, Redundant, User-Controlled Storage

Model: LOCKSS + IPFS

For Personal Data:

  • Self-hosting: NAS (Synology, QNAP) or VPS (DigitalOcean, Linode)
  • Distributed backup: Syncthing (P2P sync), Restic (encrypted backups to cloud)
  • Portable hosting: Ghost Pro, WordPress with custom domain (can migrate if provider dies)

For Public Archives:

  • IPFS (InterPlanetary File System): Content-addressed, distributed storage
  • BitTorrent: Proven P2P distribution (Archive Team uses this)
  • LOCKSS networks: Libraries collectively preserve (multiple institutions, redundant copies)

For Long-Term Preservation:

  • Open formats: Markdown, HTML, plain text (readable in 50 years)
  • Format migration: Periodic conversion as standards evolve
  • Emulation: Preserve original formats + software to read them

Key Technologies:

  • IPFS, Dat/Hypercore (distributed storage)
  • LOCKSS (institutional redundancy)
  • Open formats (Markdown, HTML, JSON)

Trade-offs:

  • Self-hosting requires technical skill and hardware
  • Distributed storage slower than centralized cloud
  • But: No single point of failure, no corporate control

Layer 4: Monetization (Avoiding Surveillance) §

Problem: Platforms need revenue. Advertising = surveillance. Subscriptions alone may not scale.

Third Way Solution: Hybrid Economic Models

Option 1: Direct User Payment

  • Subscriptions (Ghost, Fastmail, Proton)
  • One-time purchases (Obsidian, Things)
  • Donations (Wikipedia, Internet Archive)

Option 2: Cooperative Ownership

  • Users own platform collectively (Stocksy for photographers, Resonate for musicians)
  • Profits distributed to member-owners
  • Democratic governance

Option 3: Public Funding

  • Government grants (NEH, Mellon, Mozilla Foundation)
  • Public broadcasting model (NPR, BBC—funded by public, no ads)
  • University/library hosting (LOCKSS networks)

Option 4: Open Core

  • Core software free/open-source (WordPress, Ghost, Mastodon)
  • Hosting/support/premium features paid (WordPress.com, Ghost Pro)
  • Cannot enclose the core (GPL prevents proprietary forks)

Option 5: Solidarity Economy

  • Cross-subsidization (profitable projects fund loss-leaders)
  • Sliding scale (wealthy users pay more, subsidize free tiers)
  • Example: Means-based pricing (Patreon alternative)

Key Insight: No single model works for all. Need ecosystem of models, all non-surveillance.

Trade-offs:

  • Direct payment excludes those who can’t pay (need solidarity mechanisms)
  • Public funding vulnerable to political shifts
  • Cooperatives hard to scale (governance complexity)
  • But: All preferable to surveillance capitalism

Layer 5: Governance §

Problem: Platforms are dictatorships (even benevolent ones eventually betray users).

Third Way Solution: Federated, Democratic Governance

Model: Mastodon’s Federation + Co-op Governance

Federated Moderation:

  • Each instance sets own rules (no universal ToS)
  • Instances can defederate (block other instances)
  • Users choose instance that matches their values
  • If admin becomes tyrant, users migrate (account portability)

Cooperative Governance:

  • Platform owned by users/workers (one member, one vote)
  • Major decisions require supermajority (75%+ approval)
  • Transparent financials, public board meetings
  • Cannot sell to corporation (bylaws prevent acquisition)

Open Source + Forking:

  • Code is public (GPL/AGPL license)
  • If maintainers sell out, community forks (Nextcloud forked from ownCloud)
  • Prevents capture

Key Technologies:

  • ActivityPub (enables federation)
  • Cooperative bylaws (legal structure)
  • Open source licenses (GPL, AGPL)

Trade-offs:

  • Federation creates fragmentation (different instances, different rules)
  • Democratic governance is slow (voting takes time)
  • But: No single point of failure, no dictator risk

Part III: What the Third Way Looks Like in Practice §

Let’s imagine a day in the life of a Third Way internet user in 2035:

Morning: Reading and Writing §

7:00 AM — Wake up, check RSS reader (no algorithm, just chronological feeds from blogs/sites you chose)

7:30 AM — Write blog post on your site (yourname.com). Auto-syndicates to:

  • Fediverse (ActivityPub)
  • Email newsletter (subscribers you own)
  • RSS (anyone can subscribe)

All from your domain. If your hosting provider dies, you migrate (same domain, same URLs).

8:00 AM — Read replies via Webmentions (other blogs responding to yours, comments appear on your site, no centralized comment system)

Midday: Communication §

12:00 PM — Video call with friend using Jitsi (open source, self-hosted, E2E encrypted, no Zoom spying)

1:00 PM — Check Matrix (federated chat). Messages from friends on different servers (some self-hosted, some using hosting services, all interoperate)

2:00 PM — Browse Fediverse (Mastodon, Pixelfed, PeerTube). See posts from across federated instances. No ads, no algorithmic manipulation, chronological.

Evening: Entertainment and Community §

6:00 PM — Watch video on PeerTube (federated YouTube alternative, creator-owned)

7:00 PM — Listen to music on Bandcamp (artists get 82% of revenue, you own MP3s, DRM-free)

8:00 PM — Participate in forum (self-hosted Discourse, community-owned, full export available)

Night: Preservation §

10:00 PM — Automatic backup runs:

  • Your blog: Synced to NAS (RAID, redundant)
  • Photos: Syncthing to friend’s server (mutual backup)
  • Notes: Obsidian vault (Markdown files, local + cloud backup)

If any service shuts down tomorrow, you have:

  • All your data (multiple copies)
  • Your domain (persistent identity)
  • Your social graph (portable followers via ActivityPub)

You are sovereign.


Part IV: The Transition Strategy — How We Get There §

The Third Way doesn’t happen overnight. How do we transition from Platform Feudalism to Digital Sovereignty?

Phase 1: Build Alternatives (Now - 5 years) §

Goal: Prove alternatives can work at scale.

Actions:

  • Grow Mastodon/Fediverse: 10M+ users (demonstrate federation viability)
  • Launch platform co-ops: Stocksy-style models for social media, hosting, storage
  • Expand public infrastructure: Library-hosted Mastodon instances, university archives
  • Create easy on-ramps: Tools like Yunohost (one-click self-hosting), Pika (easy static sites)

Success Metrics:

  • 5% of social media users on federated platforms
  • 10+ viable platform cooperatives (profitable, member-owned)
  • 100+ universities/libraries hosting instances
  • Open-source alternatives exist for all major platforms (social, messaging, storage, video)

Phase 2: Policy Wins (5-10 years) §

Goal: Legal frameworks that enable Third Way, constrain platforms.

Actions:

  • Interoperability mandates: EU Digital Markets Act model (platforms must allow third-party clients)
  • Right to archive: Laws allowing libraries/archives to preserve copyrighted content
  • Data portability: GDPR-style requirements (full exports in usable formats)
  • Anti-monopoly enforcement: Break up Big Tech, prevent acquisitions that consolidate power

Success Metrics:

  • US/EU laws require platform interoperability
  • Copyright exceptions for preservation (fair use expanded)
  • Surveillance capitalism regulated (behavioral targeting restricted)
  • No new platform monopolies (mergers blocked)

Phase 3: Cultural Shift (10-20 years) §

Goal: Sovereignty becomes expectation, not exception.

Actions:

  • Digital literacy: Schools teach domain ownership, data sovereignty, federation
  • Cultural normalization: “Where’s your domain?” becomes as common as “What’s your email?”
  • Professional requirement: Journalists, academics, professionals expected to have sovereign presence
  • Platform stigma: Using corporate platforms seen as irresponsible (like smoking—stigmatized, not illegal)

Success Metrics:

  • 50% of internet users own domains
  • 25% of social media on federated platforms
  • Surveillance-based platforms in decline (losing users, not growing)
  • “Digital sovereignty” taught in schools

Phase 4: Infrastructure Maturity (20-30 years) §

Goal: Third Way is default, feudalism is legacy.

Actions:

  • Public infrastructure: Governments run federated instances (like public libraries run physical space)
  • Cooperative economy: Platform co-ops dominant in hosting, social media, cloud storage
  • Preservation embedded: All systems designed for 50+ year persistence
  • No more platform murders: Culture persists because infrastructure is distributed and community-owned

Success Metrics:

  • Majority of internet users on sovereign infrastructure
  • Corporate platforms either reformed (co-ops) or dead
  • Cultural memory preserved (no more GeoCities-scale losses)
  • Next generation can’t imagine Platform Feudalism (it’s history)

Part V: Objections and Responses §

Objection 1: “This is too technical for normal people” §

Response:

  • Email was “too technical” in 1995. Now everyone has email.
  • Complexity can be hidden (Ghost makes custom domains easy, Mastodon hosts handle technical bits)
  • Trade-off: Sovereignty requires some effort, but tools can minimize it

Counter-Question: Is it really “easier” to have your identity revoked, data deleted, and memories erased by platforms?

Objection 2: “Federation fragments communities” §

Response:

  • Email is federated. Do you feel “fragmented” from Gmail users if you use Fastmail? No.
  • Federation enables choice (pick instance that matches your values)
  • Interoperability prevents fragmentation (ActivityPub lets instances communicate)

Counter-Question: Isn’t platform monopoly worse fragmentation? (Twitter vs. TikTok vs. Instagram—all walled gardens)

Objection 3: “People prefer convenience over sovereignty” §

Response:

  • True in short term. But platforms eventually betray convenience (Twitter’s chaos, Facebook’s privacy violations)
  • Once betrayed, users seek alternatives (see: Twitter → Mastodon migration)
  • Convenience is temporary; sovereignty is permanent

Counter-Question: Is it convenient when the platform shuts down and you lose everything?

Objection 4: “Who will moderate a distributed internet?” §

Response:

  • Federated moderation: Each instance sets rules, defederates bad actors
  • Harder than centralized, yes. But centralized moderation has failed (harassment, hate speech, manipulation persist)
  • Trade-off: Imperfect distributed moderation > failed centralized moderation

Counter-Question: Has centralized moderation worked? (No—Facebook/Twitter full of toxicity despite armies of moderators)

Objection 5: “This requires trusting strangers to run servers” §

Response:

  • You already trust strangers (Google, Meta engineers you’ve never met)
  • Federation distributes trust (if one admin is bad, you migrate)
  • Can self-host if you want ultimate control

Counter-Question: Is trusting a for-profit corporation safer than trusting a community-run instance?

Objection 6: “Big Tech will crush alternatives” §

Response:

  • They’ll try. But open protocols are hard to kill (email survived, BitTorrent survived)
  • Network effects work both ways (once federated platforms hit critical mass, they grow)
  • Laws can help (interoperability mandates prevent lock-in)

Counter-Question: If we don’t try, Big Tech wins by default. Is surrender preferable?


Part VI: The Archaeobytologist’s Role in the Third Way §

As Archaeobytologists, what’s our work in forging the Third Way?

Role 1: Preserve the Evidence §

Archive platform murders to document what went wrong:

  • GeoCities, Vine, Google+, Tumblr NSFW purge
  • Build “Museum of Murdered Platforms” (physical/digital)
  • Use archives to teach: “This is what happens when you don’t own your ground”

Purpose: Historical memory. Can’t build future if we forget past.

Role 2: Build the Alternatives §

Forge tools and institutions that embody Three Pillars:

  • Launch preservation co-ops (community-owned archives)
  • Create sovereignty tools (easy domain setup, federated hosting)
  • Design long-term institutions (50-year orgs, LOCKSS networks)

Purpose: Demonstrate alternatives are viable. Proof of concept.

Role 3: Teach Sovereignty §

Educate next generation on digital rights and responsibilities:

  • University courses in Archaeobytology (this textbook)
  • Workshops for communities (how to own your domain, export data)
  • Public talks (TED, podcasts, op-eds)

Purpose: Cultural shift. People can’t demand sovereignty if they don’t know it exists.

Role 4: Advocate for Policy §

Fight for laws that enable Third Way:

  • Testify at hearings (right to archive, interoperability, data portability)
  • Draft model legislation (work with EFF, Creative Commons)
  • Build coalitions (libraries, journalists, activists, academics)

Purpose: Legal infrastructure. Alternatives need policy support to compete with monopolies.

Role 5: Document and Theorize §

Publish research on platform power, preservation methods, sovereignty design:

  • Academic journals (Journal of Archaeobytology, DH journals, STS venues)
  • Books (popular and scholarly)
  • Open documentation (wikis, tutorials, case studies)

Purpose: Knowledge infrastructure. Field needs canon, methods, theory.

The Complete Archaeobytologist §

You are:

  • Archivist (preserving murdered platforms)
  • Builder (forging sovereign alternatives)
  • Teacher (spreading digital literacy)
  • Advocate (fighting for policy change)
  • Scholar (documenting and theorizing)

The Third Way requires all five roles. You don’t have to do everything, but the field collectively must.


Part VII: The Archaeobytologist’s Manifesto §

We Believe: §

1. Digital culture is worth preserving.

  • Every GeoCities homepage, every Vine, every forum post—these are artifacts of human creativity and connection.
  • Platforms murder culture. We refuse to accept this.

2. Users deserve sovereignty.

  • You should own your identity, control your connections, possess your ground.
  • Platforms are landlords. We advocate for ownership.

3. Surveillance capitalism is illegitimate.

  • Monetizing behavior through tracking is exploitation.
  • We build economic models that don’t require surveillance.

4. Preservation is a moral imperative.

  • Future generations deserve access to our digital culture.
  • We are custodians, not just consumers.

5. The Third Way is possible.

  • Federated, cooperative, community-owned infrastructure can work.
  • We have the technology. We need the will.

We Commit To: §

1. Archive what platforms murder.

  • Scrape dying platforms.
  • Curate rescued artifacts.
  • Make archives accessible.

2. Build alternatives that resist murder.

  • Design for sovereignty (Three Pillars).
  • Create institutions that last 50+ years.
  • Open-source everything.

3. Teach digital sovereignty.

  • Write, speak, teach.
  • Make sovereignty accessible.
  • Raise generation that demands ownership.

4. Advocate for systemic change.

  • Fight for right to archive.
  • Demand platform interoperability.
  • Break monopolies.

5. Practice what we preach.

  • Own our domains.
  • Use federated platforms.
  • Preserve our own data.

We Reject: §

1. Platform feudalism (users as tenants)

2. Surveillance capitalism (behavior as commodity)

3. Planned obsolescence (culture murdered for profit)

4. Forced amnesia (deletion of digital history)

5. Learned helplessness (“Platforms will always win”)

We Declare: §

Archaeobytology isn’t just a discipline—it’s a movement.

We are scholars and smiths, archivists and advocates, mourners and builders.

We study the dead to prevent future murders.

We preserve the past to forge the future.

We are the Third Way.

And we are just beginning.


Conclusion: Build Something That Outlasts You §

This textbook began with a question: What is Archaeobytology?

Now you know:

  • Theory (Taxonomy, Three Pillars, Triage, Discipline Formation)
  • Methods (Excavation, Forensics, Workflow)
  • Practice (Institution Building, Sovereignty Design, Commons Governance, Memory Institutions)
  • Strategy (Political Economy, Movement Building, Public Scholarship)

You have the tools. Now the question is: What will you do?

Will you:

  • Archive a dying platform before it vanishes?
  • Build a tool that embodies sovereignty?
  • Teach a course that trains the next generation?
  • Write an op-ed that shifts public discourse?
  • Found an organization that outlasts you?

Archaeobytology doesn’t exist yet—not fully. There are no departments, no tenure-track jobs, no professional society. But there could be, if we build them.

In 20 years, this could be a recognized discipline. Students could major in it. Governments could fund it. Culture could be preserved, not murdered.

Or: This could be a footnote. A quirky experiment by scattered practitioners. Forgotten when platforms finally consolidate into permanent monopolies.

That choice is ours.

Every time you:

  • Preserve an artifact, you’re voting for the Third Way
  • Build a tool, you’re forging alternatives
  • Teach sovereignty, you’re spreading the movement
  • Advocate for policy, you’re shifting power
  • Call yourself an Archaeobytologist, you’re making the discipline real

This textbook is a beginning, not an ending. It codifies existing practice and proposes a future. But books don’t build disciplines—people do.

You, reading this now, are part of the founding generation. The choices you make—what you preserve, what you build, what you teach—will shape whether Archaeobytology becomes real.

So ask yourself:

What will you build that outlasts you?

Not what will you consume, what will you scroll, what will you post into the void of platforms that will delete it when you stop being profitable.

What will you build that future generations can find, study, and build upon?

  • A website on your own domain that persists for decades?
  • An archive of a community that would otherwise be forgotten?
  • A tool that helps others own their digital lives?
  • A course that trains students to become Archaeobytologists?
  • An institution—a journal, a conference, a center—that becomes infrastructure?

The Third Way requires builders.

Not just theorists. Not just critics. Builders.

People who preserve, create, organize, teach, and advocate.

People who look at murdered platforms and say: Never again.

People who look at surveillance capitalism and say: Not us.

People who look at the choice between feudalism and sovereignty and say: We choose the Third Way.


Final Exercise: Your Third Way Project §

Design your contribution to the Third Way. Choose one:

Option A: Preservation Project §

  • Pick a vulnerable platform
  • Design complete preservation strategy
  • Execute (or outline execution plan if resources lacking)

Option B: Sovereignty Tool §

  • Identify a sovereignty gap (something users can’t easily do)
  • Design tool that fills gap
  • Build prototype or spec for others to build

Option C: Institution §

  • Design organization that embodies Three Pillars
  • Complete business plan (funding, governance, sustainability)
  • Launch (or create plan for launch)

Option D: Movement Campaign §

  • Identify policy change needed for Third Way
  • Design 5-year campaign to achieve it
  • Begin execution (write op-ed, contact legislators, build coalition)

Option E: Pedagogical Project §

  • Design course, workshop, or curriculum
  • Create materials (syllabus, readings, assignments)
  • Teach it (or find someone who will)

Requirements (3,000+ words):

  1. Problem diagnosis (what’s broken now?)
  2. Third Way solution (how does your project fix it?)
  3. Implementation plan (concrete steps, timeline, resources)
  4. Three Pillars assessment (does it embody sovereignty?)
  5. Sustainability (how does it last 10+ years?)
  6. Impact metrics (how do you measure success?)

Then: Actually do it.

Don’t just write the plan. Execute.

Build something.

Preserve something.

Teach someone.

Advocate somewhere.

Make Archaeobytology real.

Because the Third Way doesn’t forge itself.

You forge it.

Now go.

Build something that outlasts you.


Further Reading: The Complete Archaeobytology Canon §

This textbook has cited hundreds of sources. Here’s the essential reading list—the books every Archaeobytologist should read.

Foundational Theory (Start Here) §

  1. Lessig, Lawrence. Code: Version 2.0. Basic Books, 2006.
  2. How digital architecture embodies values

  3. Zuboff, Shoshana. The Age of Surveillance Capitalism. PublicAffairs, 2019.

  4. Definitive critique of platform economics

  5. Ostrom, Elinor. Governing the Commons. Cambridge, 1990.

  6. How to manage shared resources without state or market

  7. Doctorow, Cory. The Internet Con: How to Seize the Means of Computation. Verso, 2023.

  8. Practical vision for interoperability and user power

  9. Kirschenbaum, Matthew. Mechanisms: New Media and the Forensic Imagination. MIT Press, 2008.

  10. Foundational text on digital materiality

Digital Preservation §

  1. Chun, Wendy Hui Kyong. Programmed Visions: Software and Memory. MIT Press, 2011.

  2. Ernst, Wolfgang. Digital Memory and the Archive. Minnesota, 2013.

  3. Brügger, Niels, and Ralph Schroeder, eds. The Web as History. UCL Press, 2017.

Platform Critique §

  1. Gillespie, Tarleton. Custodians of the Internet. Yale, 2018.

  2. Noble, Safiya Umoja. Algorithms of Oppression. NYU Press, 2018.

  3. Pasquale, Frank. The Black Box Society. Harvard, 2015.

Commons and Cooperation §

  1. Benkler, Yochai. The Wealth of Networks. Yale, 2006.

  2. Bollier, David. Think Like a Commoner. New Society, 2014.

  3. Scholz, Trebor. Platform Cooperativism. Rosa Luxemburg Stiftung, 2016.

Privacy and Sovereignty §

  1. Schneier, Bruce. Data and Goliath. Norton, 2015.

  2. Véliz, Carissa. Privacy Is Power. Melville House, 2020.

  3. Rushkoff, Douglas. Throwing Rocks at the Google Bus. Portfolio, 2016.

Craft and Making §

  1. Sennett, Richard. The Craftsman. Yale, 2008.

  2. Pye, David. The Nature and Art of Workmanship. Cambridge, 1968.

Archives and Memory §

  1. Derrida, Jacques. Archive Fever. Chicago, 1996.

  2. Caswell, Michelle. Urgent Archives. Routledge, 2021.

Discipline Formation §

  1. Klein, Julie Thompson. Interdisciplining Digital Humanities. Michigan, 2015.

  2. Kuhn, Thomas. The Structure of Scientific Revolutions. Chicago, 1962.

Primary Sources (Must-Read Essays) §

  1. Kahle, Brewster. “Preserving the Internet.” Scientific American, 1997.

  2. Bush, Vannevar. “As We May Think.” The Atlantic, 1945.

  3. Raymond, Eric. “The Cathedral and the Bazaar.” 1997.


The End—And The Beginning §

You’ve reached the end of this textbook.

But this is not the end of Archaeobytology.

It’s the beginning.

The field exists because you make it real.

Every artifact you preserve. Every tool you build. Every course you teach. Every policy you advocate for.

That’s Archaeobytology.

Welcome to the discipline.

Now go forth and forge the Third Way.


End of Textbook


Appendices §

The following appendices provide practical resources for Archaeobytologists:

  • Appendix A: Glossary of Terms
  • Appendix B: Essential Tools & Resources
  • Appendix C: Sample Syllabi (101, 200, 300 levels)
  • Appendix D: Teaching Resources
  • Appendix E: Professional Resources (Career Pathways, Job Descriptions, Certification)

[Appendices would be developed separately as standalone documents]


About This Textbook §

Archaeobytology: Theory and Practice of Digital Sovereignty

Author: [To be determined—likely community-authored/edited given the discipline’s nascent state]

Publication Model: Open Access

  • Free PDF download
  • Print-on-demand (estimated $40 paperback)
  • CC BY-SA 4.0 License (share, adapt, but credit and keep open)

Suggested Citation:

Archaeobytology: Theory and Practice of Digital Sovereignty. [Publisher], [Year]. [URL].

Companion Website: archaeobytology.org

  • Video lectures (18 chapters × 20 min)
  • Discussion forums
  • Tools repository
  • Syllabi database
  • Community directory

For Instructors: Instructor’s Guide available at archaeobytology.org/teaching

  • Lecture slides
  • Assignment rubrics
  • Discussion prompts
  • Quiz/exam questions

Contact: archaeobytology@[domain] for corrections, suggestions, course adoption inquiries


The textbook you hold is a founding document. By reading it, teaching from it, building on it, and critiquing it, you’re helping create a discipline.

Thank you for being part of the founding generation of Archaeobytology.

Now go build something that outlasts you.

Apparatus • Reference, Syllabi & Curricular Toolkit

Appendix A: Glossary of Terms

13 min read 2,745 words

Core Concepts §

Archaeobytology The study and practice of excavating, preserving, interpreting, and building with digital artifacts—particularly those murdered by platform shutdowns or rendered obsolete by technological change. Combines retrospective preservation (the Archive) with prospective creation (the Anvil).

Archaeobyte A digital artifact that was once alive (accessible, functional), died through platform shutdown or obsolescence, and has been preserved in some form. Exists in liminal state between death and potential resurrection. Example: GeoCities pages saved by Archive Team.

Vivibyte A digital artifact that is currently alive (accessible, functional) but exists on vulnerable infrastructure facing existential threats. The “living endangered species” of digital culture. Example: Content on Twitter/X during ownership instability.

Umbrabyte A digital artifact that is technically dead (inaccessible, non-functional) but has not been properly preserved. Exists in fragmentary or corrupted form, haunting the present through memory and partial remnants. Includes several subtypes of liminal artifacts:

  • Zombyte (formerly Necrobyte): An artifact that was dead but has been “resurrected” through external emulation or reconstruction, giving it an “undead” functionality not native to the current ecosystem. Example: Flash games running via Ruffle.
  • Xenobyte: An artifact so old or alien (orphaned code, lost encryption keys) that it is unintelligible without extensive interpretation or translation. It is the artifact on the verge of becoming permanently opaque.

Nullibyte A digital artifact known or believed to have existed but which currently resides beyond the horizon of recoverability. It is not a file; it is a “missing persons report.” Example: The 50 million songs lost in the MySpace server migration.

Cryptobyte A “digital cryptid”—an artifact rumored to exist but never verified by forensic evidence. It exists in folklore rather than the file system. Example: The “Polybius” arcade game, legendary “lost” cuts of films.

Petribyte A digital artifact so old that its original context is historical, has been durably preserved by institutions, and is treated as cultural heritage. Has achieved monumental stability. Example: ARPANET documentation preserved by Computer History Museum.


The Three Pillars of Digital Sovereignty §

Declaration (I Am) The principle that you should be able to declare your identity and existence without permission from platforms or intermediaries. Includes self-owned identity ([email protected]), persistent presence, and uncensorable voice.

Connection (Instant Message) The principle that you should be able to communicate directly with others without platform mediation, monitoring, or monetization. Includes peer-to-peer communication, portable relationships, and intentional discovery.

Ground (Digital Real Estate) The principle that you should own the infrastructure your digital life is built on, not rent it from landlords who can evict you. Includes data ownership, infrastructure control, and persistence independent of platform survival.

Digital Sovereignty The ability to exist, communicate, and build in digital space without corporate gatekeeping. Achieved through embodying all Three Pillars. Not absolute freedom (legal and social accountability remain), but freedom from arbitrary platform power.


The Archive and the Anvil §

The Archive The retrospective practice of Archaeobytology: excavating endangered artifacts, preserving them with technical and cultural fidelity, curating collections, interpreting for future generations, and providing access. Looks backward to save what’s endangered.

The Anvil The prospective practice of Archaeobytology: forging tools, protocols, and institutions that embody digital sovereignty and resist the forces that murdered previous platforms. Looks forward to build alternatives. Named for the blacksmith’s anvil where new things are forged.

Dual Soul The integration of Archive and Anvil as complementary practices. Neither is sufficient alone: Archives without alternatives accept defeat; building without remembering repeats mistakes. The complete Archaeobytologist embodies both.

The Architecture of the Archive The internal structural metaphors for organizing preserved artifacts:

  • The Seed Bank: The repository for Vivibytes. Its function is replanting; storing resilient, living artifacts (like HTML or MP3s) to prove that durable technology is possible.
  • The Haunted Forest: The repository for Umbrabytes. Its function is warning; storing the “ghosts” of murdered platforms to document what is lost when ecosystems die.
  • The Blueprint Vault: The repository for Petribytes. Its function is instruction; storing “fossils of function” (like the Away Message) as design patterns for future builders.

Preservation and Triage §

Triage The methodology for deciding what to preserve when you cannot save everything. Borrowed from emergency medicine. Requires making difficult choices about cultural significance, technical fragility, rescue feasibility, redundancy, and ethics.

The Custodial Filter Five-question ethical framework for triage decisions: (1) Cultural Significance—does this represent something that would otherwise be lost? (2) Technical Fragility—how close to disappearance? (3) Rescue Difficulty—how hard to preserve? (4) Existing Redundancy—is someone else saving this? (5) Consent and Ethics—should we preserve this?

Custodial Responsibility The ethical burden of preservation: by choosing what to save, you decide what future generations can know about the past. Every preservation decision is also a decision to let something else die. Carries weight of gatekeeping historical memory.

Triage Matrix A decision-making tool used during triage to score potential targets based on value vs. risk/effort. Helps objectify the difficult choices of what to save and what to leave.

Go/No-Go Decision The binary decision point in a preservation workflow where a team commits to a rescue operation or abandons the target. Often made under time pressure during a “War Room” scenario.

War Room The coordinated digital or physical space where a preservation team gathers during an emergency rescue (e.g., the 30 days before a site shutdown) to manage tasks, scripts, and storage in real-time.

Breadth-First Archiving A capture strategy prioritizing the top-level pages of many sites to create a “skeleton” of the web, versus Depth-First, which captures every asset of a single site. Useful when time is limited.

Platform Murder Deliberate erasure of digital artifacts by platforms through shutdown, terms of service purges, or acquisition-and-closure. Distinguished from passive obsolescence (technological decay) or neglect (link rot). Active corporate choice to kill content.


Forensic Methodology §

Forensic Materiality The concept that digital objects have a physical reality (inscriptions on a disk, voltage in memory) that can be studied as trace evidence, distinct from their symbolic meaning.

Formal Materiality The symbolic structure of digital objects (file formats, headers, code) that dictates how they behave and interact with software.

Frictional Data The “glitch” or resistance in a digital file that reveals its material history and the constraints of the medium (e.g., compression artifacts in a JPEG, corrupted headers).

Chain of Custody The documentation of the chronological history of the evidence (digital artifact). Essential in forensics to prove that the data analyzed is the same data originally collected.

Magic Numbers Unique sequences of bytes at the beginning of a file that identify its format. Used in forensics to identify file types even if extensions are missing or renamed.

Forensic Image A bit-for-bit copy of a storage media (hard drive, floppy disk). Unlike a standard file copy, it captures deleted files, slack space, and system data essential for recovery.


Technical Concepts §

Web Scraping Automated extraction of data from websites using tools like wget, HTTrack, or custom scripts. Can range from simple HTML downloads to complex JavaScript rendering. Often operates in legal gray area when done without platform permission.

API Harvesting Using a platform’s Application Programming Interface to bulk-download content. More reliable than scraping when available, but platforms control API access and can revoke it.

Emulation Running old software or systems in a simulated environment. Allows obsolete programs (Flash games, DOS applications) to function on modern hardware. Preserves not just files but user experience. Example: Ruffle emulator.

Emulation-as-Service The delivery of emulation via a web browser, allowing users to interact with obsolete software without installing local emulators. The Internet Archive’s DOSBox implementation is a prime example.

Fidelity Ladder The spectrum of preservation quality: Level 1 (Documentation/Screenshots) → Level 2 (Static Archive) → Level 3 (Emulation) → Level 4 (Resurrection/Rebuilt Backend).

Format Migration Converting files from obsolete formats to current standards to ensure long-term accessibility. Risk: May lose fidelity or functionality in translation.

Bit Rot Gradual degradation of digital storage media over time. Hard drives fail, CDs deteriorate, flash memory loses charge. Requires active preservation through redundant copies and periodic data migration.

Link Rot The phenomenon of hyperlinks breaking over time as the pages they point to are moved or deleted. A primary driver of the “vanishing web.”

Dark Archive A collection of preserved material that is not accessible to the public, often due to copyright, privacy, or donor restrictions. Preserved for the future “when the copyright expires” or for authorized researchers.

The 3-2-1 Rule The standard for data redundancy: 3 copies of data, on 2 different media types, with 1 copy off-site. The baseline for avoiding data loss.

WARC (Web ARChive format) ISO standard format for archiving web content. Stores HTTP headers, request/response data, and metadata. Used by Internet Archive’s Wayback Machine. Preserves not just content but context.

LOCKSS (Lots of Copies Keep Stuff Safe) Distributed digital preservation system and philosophy. Multiple institutions maintain copies of collections; if one fails, others survive. Embodies redundancy principle.

Metadata “Data about data”—information describing an artifact’s context, provenance, technical characteristics, and relationships. Essential for making preserved artifacts discoverable and interpretable.


Institutional and Economic Terms §

The Archive Business Model Organizational design for sustainable preservation. Includes funding sources (grants, donations, subscriptions, services), governance structure (non-profit, cooperative, hybrid), and technical infrastructure. Must survive 50+ years to succeed.

The Anvil Business Model (The Foundry) Organizational design for profitable sovereignty tools that don’t become extractive platforms. Includes revenue models that avoid surveillance capitalism. Must embody Three Pillars in business design itself.

Heroic Founder Problem The organizational vulnerability where a project relies entirely on the energy, resources, or knowledge of a single individual. If the founder burns out or leaves, the project dies.

Federated Architecture System design where multiple independent servers (instances) interoperate using open protocols. No central authority controls the network. Example: Mastodon.

Platform Capitalism Economic system where digital platforms extract value by controlling access to networks, users, and data. Creates walled gardens, lock-in effects, and surveillance business models.

Surveillance Capitalism Business model based on extracting behavioral data as raw material for prediction products sold to advertisers. Platforms surveil users to monetize attention. Incompatible with digital sovereignty.

Enshittification The lifecycle of platform decay where services first offer value to users to lock them in, then abuse users to capture business customers, and finally abuse both to capture value for shareholders.

Platform Feudalism An economic arrangement where users act as “tenant farmers” on digital land owned by platforms, creating content and value without owning the “ground” or having rights to the infrastructure.

Adversarial Interoperability The ability to create a new tool that plugs into an existing one without the permission of the original tool’s maker. A key strategy for reclaiming digital sovereignty (coined by Cory Doctorow).

Commons Governance Elinor Ostrom’s framework for collectively managing shared resources. Applied to digital preservation through principles like Graduated Sanctions (rule violations met with increasing penalties rather than immediate expulsion).

Open Core A business model where the core software is open source and free, but advanced features or hosting are paid. A common model for sovereign tech businesses (“Foundries”).

Exit to Community (E2C) A strategy for transferring ownership of a platform or company from investors/founders to its user community, often through a cooperative model or trust.

The Third Way A digital ecosystem that rejects both corporate centralization (Big Tech/Feudalism) and unmanaged chaos, prioritizing sovereignty, federation, and commons governance. Not a utopia, but a necessary alternative.

Exit Rights The technical and legal ability to leave a platform without losing your data, social connections, or identity. A prerequisite for sovereignty.

Pluralism The coexistence of multiple ownership models (state, corporate, cooperative, personal) to ensure systemic resilience.


Right to Be Forgotten Legal concept that individuals can request deletion of personal data. Creates tension with preservation: historians want to save everything, but privacy advocates prioritize consent and erasure.

Fair Use / Fair Dealing Legal doctrine allowing limited use of copyrighted material without permission for purposes like criticism, education, research, and preservation.

Context Collapse When content created for one audience becomes visible to a different audience. Common in archives when private/semi-private content is preserved and made accessible.

Informed Consent Ethical principle that people should understand and agree to how their data/content is used. Complicated in preservation where users often didn’t expect permanent archiving.

Custodial Ethics Framework for responsible stewardship of preserved artifacts, prioritizing harm reduction and transparency.


Movement and Discipline Terms §

Discipline Formation Process by which scattered practices become recognized academic/professional field. Requires intellectual coherence, institutional infrastructure, and external recognition.

Boundary Work Defining a discipline by exclusion—stating what it is NOT. Clarifies distinct identity.

Knowledge Infrastructure Journals, conferences, textbooks, handbooks, etc., that standardize and disseminate a field’s knowledge.

Institutional Anchors Universities, centers, institutes, labs, and programs that provide stable homes for a discipline.

Professional Pathways Clear career routes for people trained in a discipline. Essential for field sustainability.

Movement Building Strategic work to grow discipline from scattered practice to recognized field.

Trading Zone A space where different disciplines (e.g., Computer Science and History) can collaborate using a shared “pidgin language” without merging completely. Essential for the coalition model of Archaeobytology.

The “Gladwell Moment” The point where a complex academic field gains mainstream visibility through a popular book or media event. A milestone in public visibility.

The Tenure Trap The academic risk where public scholarship (op-eds, advocacy) is undervalued by promotion committees, disincentivizing engagement.


Historical Platforms and Projects §

GeoCities Web hosting service (1994-2009) that gave millions of people free homepages. Canonical example of platform murder.

Vine Short-form video platform (2012-2017) known for 6-second loops. Example of cultural significance vs. preservation difficulty.

Flash Player Multimedia platform by Adobe (1996-2020). Example of a tech ecosystem death that created millions of Umbrabytes.

Internet Archive Non-profit digital library (1996-present). Gold standard for institutional preservation.

Archive Team Guerrilla digital archiving collective (2009-present). Fast, agile, preserves dying platforms.

Mastodon Federated social network (2016-present). Example of sovereign, federated architecture in practice.


Digital Humanities Field using computational methods for humanities research. Related to but distinct from Archaeobytology.

Media Archaeology Theoretical field excavating dead media. Provides theoretical foundation.

Library and Information Science (LIS) Professional field managing information collections. Provides standards and ethics.

Science and Technology Studies (STS) Field studying science/technology and society. Provides frameworks for power and politics.

Platform Studies Examining how platforms shape cultural production.


Key Thinkers and Works §

Brewster Kahle Founder of Internet Archive. Builder-Evangelist.

Cory Doctorow Activist, author. Advocate for adversarial interoperability.

Elinor Ostrom Nobel laureate. Theorist of commons governance.

Shoshana Zuboff Theorist of surveillance capitalism.

Lawrence Lessig Legal scholar. “Code is Law.”

Wendy Hui Kyong Chun Media theorist. “The Enduring Ephemeral.”

Matthew Kirschenbaum Digital humanities scholar. “Forensic Materiality.”


Acronyms and Abbreviations §

API — Application Programming Interface CAPTCHA — Completely Automated Public Turing test to tell Computers and Humans Apart CSS — Cascading Style Sheets DH — Digital Humanities DMCA — Digital Millennium Copyright Act DNS — Domain Name System DRM — Digital Rights Management E2E / E2EE — End-to-End Encryption EFF — Electronic Frontier Foundation ENS — Ethereum Name Service GDPR — General Data Protection Regulation HTML — HyperText Markup Language HTTP/HTTPS — HyperText Transfer Protocol (Secure) ICANN — Internet Corporation for Assigned Names and Numbers IPFS — InterPlanetary File System IRB — Institutional Review Board ISP — Internet Service Provider LIS — Library and Information Science LOC — Library of Congress NARA — National Archives and Records Administration NEH — National Endowment for the Humanities NSF — National Science Foundation P2P — Peer-to-Peer POSSE — Publish On your Own Site, Syndicate Elsewhere RSS — Really Simple Syndication SMTP — Simple Mail Transfer Protocol STS — Science and Technology Studies TOS — Terms of Service URI/URL — Uniform Resource Identifier/Locator WARC — Web ARChive format W3C — World Wide Web Consortium


Concepts from the Textbook Chapters §

Archive Sustainability Matrix Three-dimensional framework from Chapter 11 for designing preservation organizations: Funding, Governance, Technical Infrastructure.

Foundry Business Matrix Framework from Chapter 12 for building sovereign businesses: What to Sell × Revenue Model × Business Structure.

Sovereignty Stack Six-layer infrastructure analysis from Chapter 15: (1) Physical, (2) Network, (3) Identity, (4) Storage, (5) Application, (6) Economic. Used to audit ownership and control.

Movement-Building Matrix Five-dimensional framework from Chapter 16 for discipline formation: Knowledge Infrastructure, Institutional Anchors, Professional Pathways, Public Visibility, Policy Advocacy.

Public Intellectual Toolkit Five skills from Chapter 17 for translating research into impact: Writing for Audiences, Media Engagement, Public Speaking, Platform Building, Policy Influence.


End of Appendix A: Glossary of Terms

Next: Appendix B — Essential Tools & Resources

Apparatus • Reference, Syllabi & Curricular Toolkit

Appendix B: Essential Tools & Resources

9 min read 1,937 words

Introduction §

This appendix provides a curated catalog of tools, software, services, and resources essential for Archaeobytological practice. Tools are organized by function and annotated with:

  • Purpose: What the tool does
  • Skill Level: Beginner, Intermediate, Advanced
  • Cost: Free, Freemium, Paid
  • Platform: Windows, macOS, Linux, Web-based
  • Open Source: Yes/No

Tools are current as of 2025 but the digital preservation landscape evolves rapidly. Check the Archaeobytology community wiki (archaeobytology.org/wiki) for updates.


I. Web Archiving & Scraping Tools §

1. Wget §

Purpose: Command-line tool for downloading websites recursively
Skill Level: Beginner-Intermediate
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.gnu.org/software/wget/

What It Does: Downloads web pages and their linked resources (images, CSS, JavaScript). Creates mirror copies of websites on your local machine.

Basic Usage:

bash
wget --recursive --level=2 --no-parent --wait=1 https://example.com

Best For: Static HTML sites, simple scraping projects

Limitations: Doesn’t handle JavaScript-heavy sites well, can’t navigate login walls


2. HTTrack §

Purpose: Website copier with GUI interface
Skill Level: Beginner
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.httrack.com/

What It Does: Similar to wget but with graphical interface. Easier for beginners who don’t want command-line tools.

Best For: One-time website archiving, beginners

Limitations: Slower than command-line tools, less flexible configuration


3. ArchiveBox §

Purpose: Self-hosted web archiving platform
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux, Docker
Open Source: Yes
Website: https://archivebox.io/

What It Does: Creates permanent archives of web pages including HTML, screenshots, PDFs, videos, and git repositories. Provides web interface for browsing archives.

Features:

  • Multiple capture methods (wget, Chrome headless, youtube-dl, etc.)
  • Scheduled archiving (cron jobs)
  • Full-text search
  • Deduplication

Best For: Personal archiving projects, research collections, small organizations

Setup Complexity: Requires server or Docker knowledge


4. Webrecorder (now Conifer) §

Purpose: Browser-based interactive web archiving
Skill Level: Beginner
Cost: Free
Platform: Web (browser extension also available)
Open Source: Yes
Website: https://conifer.rhizome.org/

What It Does: Records your browsing session including JavaScript interactions, videos, and dynamic content. Creates WARC (Web ARChive) files you can replay.

Features:

  • Captures JavaScript-heavy sites
  • Records social media feeds (Twitter, Instagram)
  • Exports to standard WARC format
  • Replay archives offline

Best For: Social media archiving, dynamic websites, personal projects

Unique Advantage: Works in browser, no installation required


5. Heritrix §

Purpose: Industrial-strength web crawler
Skill Level: Advanced
Cost: Free
Platform: Java (cross-platform)
Open Source: Yes
Website: https://github.com/internetarchive/heritrix3

What It Does: Internet Archive’s production crawler. Designed for massive-scale archiving (billions of URLs).

Features:

  • Highly configurable crawl policies
  • Distributed crawling
  • Respects robots.txt
  • Creates WARC files

Best For: Large institutions, comprehensive web archiving

Limitations: Steep learning curve, requires significant infrastructure


6. Browsertrix Crawler §

Purpose: High-fidelity browser-based crawling
Skill Level: Intermediate-Advanced
Cost: Free
Platform: Docker
Open Source: Yes
Website: https://github.com/webrecorder/browsertrix-crawler

What It Does: Uses real browsers (Chrome) to capture JavaScript-heavy sites with perfect fidelity. Creates WARC files.

Best For: Modern web apps, single-page applications, sites requiring JavaScript


7. Archive-It §

Purpose: Subscription web archiving service
Skill Level: Beginner
Cost: Paid (subscription based on storage)
Platform: Web-based
Open Source: No
Website: https://archive-it.org/

What It Does: Managed web archiving service by Internet Archive. Point-and-click interface for creating and managing web archives.

Features:

  • Scheduled recurring crawls
  • Metadata management
  • Public or private collections
  • Integration with Wayback Machine

Best For: Institutions without technical staff, organizations needing reliable managed service

Cost: Starts ~$1,500/year for small collections


8. Wayback Machine Downloader §

Purpose: Retrieve websites from the Internet Archive
Skill Level: Intermediate
Cost: Free
Platform: Ruby (cross-platform)
Open Source: Yes
Website: https://github.com/hartator/wayback-machine-downloader

What It Does: Downloads entire websites from the Wayback Machine to your local computer.

Basic Usage:

bash
wayback_machine_downloader http://example.com

Best For: Resurrecting dead sites that were not archived locally but exist in IA.


II. Media Preservation Tools §

9. yt-dlp §

Purpose: Video downloader for YouTube and 1000+ sites
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://github.com/yt-dlp/yt-dlp

What It Does: Downloads videos from streaming platforms including metadata, subtitles, thumbnails. The active fork of youtube-dl.

Basic Usage:

bash
yt-dlp --write-description --write-info-json --write-thumbnail https://youtube.com/watch?v=VIDEO_ID

Best For: Video archiving, preserving YouTube/Vimeo/TikTok content


Purpose: Image gallery downloader
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://github.com/mikf/gallery-dl

What It Does: Downloads images from image hosting sites (Imgur, Flickr, DeviantArt, Twitter, etc.).

Best For: Image archiving, art preservation, meme collections


11. FFmpeg §

Purpose: Multimedia conversion and processing
Skill Level: Intermediate-Advanced
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://ffmpeg.org/

What It Does: The Swiss Army knife of audio/video. Converts formats, extracts frames, creates thumbnails, transcodes for preservation.

Best For: Format migration, creating preservation masters, generating access copies

Example:

bash
ffmpeg -i input.flv -c:v libx264 -c:a aac output.mp4

12. ShareX / CleanShot X §

Purpose: Advanced screen capture
Skill Level: Beginner
Cost: ShareX (Free), CleanShot (Paid)
Platform: Windows (ShareX), macOS (CleanShot)
Open Source: ShareX (Yes)
Website: https://getsharex.com/

What It Does: Captures screenshots, GIFs, and scrolling windows. Essential for documenting “ephemeral” interfaces that cannot be scraped (e.g., Snapchats, dying apps).

Best For: Documentation of UI/UX, capturing unscrapable content.


III. Emulation & Obsolescence Tools §

13. Flashpoint Archive §

Purpose: Flash game and animation preservation
Skill Level: Beginner
Cost: Free
Platform: Windows, Linux
Open Source: Partially
Website: https://bluemaxima.org/flashpoint/

What It Does: Preserves and plays 150,000+ Flash games and animations using embedded emulators.

Best For: Playing preserved Flash content, research, nostalgia


14. Ruffle §

Purpose: Flash Player emulator in Rust
Skill Level: Beginner
Cost: Free
Platform: Web (browser extension), Desktop
Open Source: Yes
Website: https://ruffle.rs/

What It Does: Open-source Flash Player replacement that runs in browsers. Critical for “resurrecting” Petribytes without the original proprietary plugin.

Best For: Viewing archived Flash content, embedding Flash in modern websites


15. MAME (Multiple Arcade Machine Emulator) §

Purpose: Arcade game preservation
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.mamedev.org/

What It Does: Emulates arcade hardware to preserve vintage arcade games.

Best For: Arcade game preservation, historical research


16. DOSBox §

Purpose: DOS emulator
Skill Level: Beginner
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.dosbox.com/

What It Does: Emulates MS-DOS environment for running old DOS games and software.

Best For: 1980s-1990s software preservation


17. RetroArch §

Purpose: Frontend for emulators
Skill Level: Intermediate
Cost: Free
Platform: Cross-platform
Open Source: Yes
Website: https://www.retroarch.com/

What It Does: A unified interface for running emulators (cores) for dozens of consoles (NES, SNES, PlayStation, etc.).

Best For: Console history preservation, gaming.


IV. Forensics & Data Recovery §

18. The Sleuth Kit (TSK) / Autopsy §

Purpose: Digital forensics and file recovery
Skill Level: Advanced
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.sleuthkit.org/

What It Does: Analyzes disk images, recovers deleted files, examines file systems. Autopsy is the GUI; TSK is the command line.

Best For: Forensic analysis of hard drives, recovering deleted content, “Digging in the Dirt” (Chapter 8).


19. FTK Imager §

Purpose: Disk imaging tool
Skill Level: Intermediate
Cost: Free
Platform: Windows
Open Source: No
Website: https://www.exterro.com/ftk-imager

What It Does: Creates forensic disk images (bit-by-bit copies) without altering the original evidence.

Best For: Creating preservation masters of physical media.


20. PhotoRec §

Purpose: File recovery tool
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.cgsecurity.org/wiki/PhotoRec

What It Does: Recovers deleted files from hard drives, memory cards, etc., by recognizing file signatures (magic numbers).

Best For: Recovering accidentally deleted content, salvaging corrupted media.


21. Hex Editors (HxD, 0xED) §

Purpose: Raw binary editing
Skill Level: Advanced
Cost: Free
Platform: Windows (HxD), macOS (0xED)
Open Source: Varies
Website: https://mh-nexus.de/en/hxd/

What It Does: View and edit the raw binary data of a file. Essential for identifying “Magic Numbers” when file extensions are missing.

Best For: Forensic analysis, fixing corrupted headers.


V. Metadata & Organization §

22. ExifTool §

Purpose: Metadata reading/writing
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://exiftool.org/

What It Does: The industry standard for reading, writing, and editing meta-information in a wide variety of files.

Best For: Extracting metadata, adding preservation info (provenance) to files.


23. DROID (Digital Record Object Identification) §

Purpose: File format identification
Skill Level: Intermediate
Cost: Free
Platform: Windows, macOS, Linux (Java)
Open Source: Yes
Website: https://digital-preservation.github.io/droid/

What It Does: Identifies file formats and versions using the PRONOM registry.

Best For: Surveying collections, format migration planning, identifying “Xenobytes.”


24. Zotero §

Purpose: Reference management
Skill Level: Beginner
Cost: Free
Platform: Windows, macOS, Linux
Open Source: Yes
Website: https://www.zotero.org/

What It Does: Organize research sources, generate citations, save snapshots of web pages.

Best For: Academic research, maintaining a personal bibliography.


25. Obsidian §

Purpose: Knowledge base
Skill Level: Beginner-Intermediate
Cost: Free for personal use
Platform: Cross-platform
Open Source: No (Modules are open)
Website: https://obsidian.md/

What It Does: A knowledge base that works on top of a local folder of plain text Markdown files.

Best For: Maintaining “sovereign” personal knowledge graphs, documentation.


VI. Sovereignty & Infrastructure (The Anvil) §

26. Ghost §

Purpose: Sovereign publishing platform
Skill Level: Beginner (hosted) to Intermediate (self-hosted)
Cost: Freemium (Ghost Pro) or Free (self-hosted)
Platform: Web-based (Node.js)
Open Source: Yes
Website: https://ghost.org/

What It Does: Open-source platform for blogging, newsletters, and memberships.

Best For: “Ground” ownership, “POSSE” publishing (Chapter 17).


27. Mastodon §

Purpose: Federated social networking server
Skill Level: Advanced (self-hosting), Beginner (using)
Cost: Free (software)
Platform: Web-based (Ruby)
Open Source: Yes
Website: https://joinmastodon.org/

What It Does: Essential for the “Connection” pillar and federated infrastructure.

Best For: Social networking without corporate control.


28. IPFS (InterPlanetary File System) §

Purpose: Distributed file storage protocol
Skill Level: Advanced
Cost: Free
Platform: Cross-platform
Open Source: Yes
Website: https://ipfs.tech/

What It Does: Content-addressed, peer-to-peer file system. Files stored across network, retrieved by hash.

Best For: “Seed Bank” model, censorship-resistant storage.


29. Tor Browser / Onion Services §

Purpose: Anonymity and censorship resistance
Skill Level: Beginner
Cost: Free
Platform: Cross-platform
Open Source: Yes
Website: https://www.torproject.org/

What It Does: Routes traffic through a distributed network to conceal location and usage.

Best For: Privacy, bypassing censorship, accessing “Dark Archives.”


30. Signal §

Purpose: Encrypted communication
Skill Level: Beginner
Cost: Free
Platform: Mobile, Desktop
Open Source: Yes
Website: https://signal.org/

What It Does: End-to-end encrypted messaging.

Best For: “War Room” coordination, secure team communication.


VII. Storage & Backup §

31. Nextcloud §

Purpose: Self-hosted cloud storage
Skill Level: Intermediate
Cost: Free (self-hosted)
Platform: Web-based (PHP)
Open Source: Yes
Website: https://nextcloud.com/

What It Does: Personal cloud storage like Dropbox but self-hosted. Sync files across devices.

Best For: “Ground” ownership, data sovereignty.


32. Syncthing §

Purpose: Peer-to-peer file synchronization
Skill Level: Beginner-Intermediate
Cost: Free
Platform: Cross-platform
Open Source: Yes
Website: https://syncthing.net/

What It Does: Syncs files between devices without central server.

Best For: Personal backups, avoiding “The Cloud.”


33. Restic §

Purpose: Encrypted backup program
Skill Level: Intermediate
Cost: Free
Platform: Cross-platform
Open Source: Yes
Website: https://restic.net/

What It Does: Fast, encrypted, deduplicated backups to local or cloud storage.

Best For: Secure “3-2-1” backups.


VII. Reference Resources §

34. Archive Team Wiki §

Website: https://wiki.archiveteam.org/ Purpose: Documentation of rescue projects (“War Rooms”), specific platform scripts, and warrior appliances.

35. PRONOM §

Website: https://www.nationalarchives.gov.uk/PRONOM Purpose: The technical registry of file formats. Used by DROID to identify files.

36. COPTR (Community Owned digital Preservation Tool Registry) §

Website: https://coptr.digipres.org/ Purpose: A wiki listing hundreds of digital preservation tools.


End of Appendix B

Apparatus • Reference, Syllabi & Curricular Toolkit

Appendix C: Sample Syllabi & Curricular Models

5 min read 959 words

Introduction §

This appendix provides ready-to-use curricula for teaching Archaeobytology in three contexts:

  1. Academic Degrees: Full-semester courses for the 101 (Theory), 200 (Methods), and 300 (Systems) levels.
  2. Professional Development: A 2-day intensive workshop for librarians and archivists.
  3. Interdisciplinary Modules: A 2-week “injection” unit for Computer Science or History departments.

Model 1: The Undergraduate Survey (ARCH 101) §

Title: Introduction to Archaeobytology: Digital Culture and the Art of Resistance
Duration: 15 Weeks
Prerequisites: None

Course Description: What happens when digital platforms die? This course introduces students to the study of “murdered” digital culture. We move beyond the “digital dualism” that separates the “real” from the “virtual” to understand digital artifacts as material objects subject to decay, destruction, and preservation. Students will learn to classify artifacts using the Archaeobyte Taxonomy and evaluate the power structures of the digital world through the Three Pillars of Sovereignty.

Learning Objectives:

  • Taxonomy: Classify artifacts as Vivibytes, Umbrabytes, or Petribytes.
  • Ethics: Apply the “Custodial Filter” to preservation decisions.
  • Theory: Analyze platforms using the Declaration/Connection/Ground framework.
  • Practice: Conduct a basic “Digital Life Audit” of personal data.

Weekly Outline:

  • Weeks 1–3: Foundations. The history of platform death (GeoCities to Vine). Defining the Archaeobyte.
  • Weeks 4–6: Theory. The Three Pillars of Sovereignty. The Archive vs. The Anvil.
  • Weeks 7–10: Excavation. Introduction to the Wayback Machine and basic scraping concepts.
  • Weeks 11–13: Ethics. Privacy, consent, and the “Right to be Forgotten.”
  • Weeks 14–15: The Future. Building sovereign alternatives.

Key Assignment: The Platform Autopsy
Students select a dead platform (e.g., Google Reader, Vine) and write a 2,000-word “Coroner’s Report” analyzing its cause of death, the community displacement, and the surviving fossil record.


Model 2: The Methodological Seminar (ARCH 200) §

Title: Digital Preservation Methods: Excavation, Forensics, and Triage
Duration: 15 Weeks (Lab/Seminar Hybrid)
Prerequisites: ARCH 101 or CS 101

Course Description: This is a hands-on technical course. Students transition from studying digital death to preventing it. We focus on the “Rescue Phase”—the critical window between a platform’s shutdown announcement and its deletion. Students will learn command-line scraping, metadata extraction, and forensic disk imaging.

Technical Stack:

  • Command Line: Wget, Youtube-dl, ffmpeg.
  • Forensics: The Sleuth Kit (TSK), ExifTool.
  • Emulation: Ruffle, DOSBox.

Weekly Outline:

  • Weeks 1–4: Reconnaissance. Mapping site architecture and identifying hidden assets.
  • Weeks 5–8: Extraction. Wget scripting, API interaction, and WARC file generation.
  • Weeks 9–12: Forensics. Recovering data from “bit-rotted” formats and corrupted media.
  • Weeks 13–15: Triage. Running “Fire Drill” simulations where students must prioritize data under time constraints.

Key Assignment: The Rescue Simulation
A 48-hour take-home exam. Students are given a “dying” test server (hosted by the instructor) scheduled to auto-delete in 48 hours. They must scrape, validate, and package the site’s content before the clock runs out.


Model 3: The Graduate Capstone (ARCH 300) §

Title: Institution Building and Strategic Infrastructure
Duration: 15 Weeks
Prerequisites: ARCH 200

Course Description: How do we build structures that last 50 years? This graduate seminar shifts from the individual practitioner to the institutional level. Students study the failures of previous preservation attempts (the “Heroic Founder” problem, funding collapses) and design sustainable organizations—Archives, Foundries, and Seed Banks—that can endure.

Learning Objectives:

  • Design: Create governance models based on Ostrom’s Principles for the Commons.
  • Economics: Develop business models for “Sovereign Foundries” (non-extractive tech).
  • Strategy: Map a 20-year “Movement Building” strategy for a specific digital community.

Module Structure:

  • Module 1: The Institutional Void. Diagnosing why current archives fail.
  • Module 2: The Business of the Archive. Designing non-profit and hybrid revenue models.
  • Module 3: The Seed Bank. Designing distributed/federated governance (LOCKSS).
  • Module 4: The Haunted Forest. Curating memory institutions for the public.
  • Module 5: The Anvil. Building tools for the post-platform future.

Key Assignment: The Institutional Prospectus
Students produce a 30-page “Launch Deck” for a new institution (e.g., “The Museum of Flash Games” or “The Decentralized Social Archive”), including bylaws, 10-year budget, and technical architecture.


Model 4: The Community Workshop (Non-Academic) §

Title: Digital Self-Defense: A Weekend Intensive
Duration: 2 Days (12 Hours)
Audience: Community archivists, activists, librarians, and the general public.
Goal: Democratize Archaeobytology skills for immediate community use.

Day 1: The Archive (Defense)

  • Morning (Theory): “Why Your Data is Disappearing.” Understanding platform terms of service and the lifecycle of data.
  • Afternoon (Practice): “The Personal Rescue.” Participants bring their own laptops and perform a “Takeout” of their data from Google/Facebook/Twitter. We teach them how to verify, store, and organize these dumps so they are readable without the platform.

Day 2: The Anvil (Offense)

  • Morning (Theory): “Sovereignty 101.” Buying a domain name, understanding DNS, and the difference between “renting” and “owning” digital ground.
  • Afternoon (Practice): “Planting the Flag.” Every participant leaves with a personal website or digital garden running on their own domain, independent of social media silos.

Model 5: The “Trojan Horse” Module (Interdisciplinary) §

Title: The Materiality of the Internet
Duration: 2 Weeks (Insertable into History, Media Studies, or CS syllabi)
Goal: Plant the seeds of Archaeobytology in established disciplines.

Week 1: Excavating the Recent Past

  • Reading: Chapter 2 (Taxonomy) and Chapter 6 (The Warning of Rented Land).
  • Activity: “Digital Stratigraphy.” Students look at a single website (e.g., the White House site) via the Wayback Machine across 10 years and map the changing layers of technology and ideology.

Week 2: The Ethics of Memory

  • Reading: Chapter 9 (The Custodial Filter).
  • Activity: “The Triage Committee.” The class is presented with a hypothetical server drive from a defunct extremist forum. They must debate and vote on whether to destroy it, seal it, or publish it, using the Custodial Filter framework.

Apparatus • Reference, Syllabi & Curricular Toolkit

Appendix D: Teaching Resources

The Instructor’s Toolkit

5 min read 1,003 words

I. Introduction §

This appendix serves as the Instructor’s Companion. It translates the theoretical concepts of the previous 18 chapters into actionable classroom mechanics. It answers the professor’s question: “How do I actually teach this?”

Philosophy: We teach Archaeobytology not just to transfer knowledge, but to train practitioners. Theoretical disputes about “digital dualism” are useful, but the ultimate goal is to produce graduates capable of saving history. Every assignment should result in a portfolio piece—a rescued artifact, a forensic report, or a strategic plan.

This toolkit provides ready-made discussion prompts, assignment sheets, rubrics, and technical lab guides to allow instructors to “plug and play” the curriculum.


II. Discussion Facilitation Guides §

Framework: These prompts are designed to move students from “gut reaction” to “systematic analysis” using the Custodial Filter (Significance, Fragility, Feasibility, Redundancy, Ethics).

Scenario 1: The Deleted Confession §

  • The Case: A beloved celebrity posts a racist tweet at 2:00 AM. Five minutes later, they delete it. No screenshots exist yet. You captured it in your automated feed scraper.
  • The Conflict: Accountability (History) vs. Right to be Forgotten (Privacy).
  • Facilitation Tip: Ask students to vote: Delete or Save? Then ask: “Does the celebrity’s public status change the ethical calculus? What if it was your 15-year-old cousin instead?”
  • Key Concept: Public Interest Exemption.

Scenario 2: The Teen’s Blog §

  • The Case: A 30-year-old professional discovers their rigorous 14-year-old coming-out blog is still online and archived by the Wayback Machine. They beg you to remove it, citing professional embarrassment. The blog is a unique primary source on queer youth culture in the 2000s.
  • The Conflict: Future History (Collective Value) vs. Present Consent (Individual Harm).
  • Facilitation Tip: Use the “Temporal Distance” argument. Does the harm fade over time? Is the 14-year-old a different legal entity than the 30-year-old?
  • Key Concept: The Right to Curate the Self.

Scenario 3: The Hate Forum §

  • The Case: A notorious white supremacist forum is shutting down. It contains evidence of radicalization pathways, but also hate speech and doxxing of victims. You have the bandwidth to mirror it.
  • The Conflict: Research Value (Understanding Extremism) vs. Harm Reduction (Amplifying Hate).
  • Facilitation Tip: Discuss “Dark Archiving.” Can we save it without publishing it? Who gets access?
  • Key Concept: The Quarantine Archive.

III. Assignment Templates §

Assignment 1: The Platform Autopsy §

  • Task: Select a dead platform (e.g., Vine, Google+, Friendster) and write a 2,000-word “Coroner’s Report.”
  • Requirements:
    1. Life History: When was it born? Who used it? (Demographics).
    2. Cause of Death: Diagnose the “Murder Weapon.” Was it corporate strategy, neglect, acquisition, or server failure?
    3. Taxonomic Analysis: Classify the surviving artifacts. Are they Petribytes (frozen)? Umbrabytes (shadowy/incomplete)?
    4. Preservation Status: Where is the body? (Internet Archive, torrents, lost forever).
  • Learning Outcome: Understanding the lifecycle of digital platforms.

Assignment 2: The 72-Hour Triage Simulation §

  • Task: You are the lead archivist. A niche fanfiction site (“FanFicX”) has announced it will shut down in 72 hours.
  • Constraints: You have 1TB of storage, 5 volunteer archivists, and limited bandwidth. The site is 50TB. You cannot save everything.
  • Deliverable: A Triage Matrix and Action Plan.
    • What do you save first? (Text? Images? Comments? Metadata?)
    • What do you abandon?
    • How do you deploy your 5 volunteers?
  • Learning Outcome: Applying the Custodial Filter under pressure.

IV. Lab Exercises §

Lab 1: The Personal Rescue §

  • Objective: Demystify the “black box” of archiving by making it personal.
  • Task: Students must archive a single page of their own digital footprint (e.g., their own Twitter profile, a personal blog post) that is not currently backed up.
  • Tools:
    • Internet Archive “Save Page Now”: For public-facing preservation.
    • Webrecorder (Conifer): For capturing dynamic content/scripts.
  • Deliverable: A link to the stable WARC file and a paragraph reflecting on the difference between the “live” site and the “captured” version.

Lab 2: The Forensic Gaze §

  • Objective: Understand “forensic materiality” by looking beneath the interface.
  • Task: Download a “corrupted” image file provided by the instructor. Open it in a Hex Editor (HxD or 0xED).
  • Instructions:
    1. Identify the file header (Magic Number).
    2. Find the text string hidden in the metadata comments.
    3. Repair the broken header to make the image viewable again.
  • Deliverable: The “Secret Message” found in the file and the repaired image.

V. Case Study Teaching Notes §

1. GeoCities Rescue (2009) §

  • Focus: Scale, Speed, and “Crisis Archiving.”
  • Teaching Point: This was the “Dunkirk Moment” of web archiving. Use this to discuss the transition from “Polite spidering” to “Guerrilla scraping.”
  • Discussion: Was it ethical to scrape personal pages without consent? (Answer: Yes, because the alternative was total annihilation).

2. Tumblr NSFW Purge (2018) §

  • Focus: Marginalized communities and “Algorithmic Eviction.”
  • Teaching Point: This demonstrates how Terms of Service changes act as “Soft Deletion.”
  • Discussion: How do definitions of “Obscenity” serve as tools for erasure? Why are queer spaces disproportionately targeted?

3. Mastodon (Present) §

  • Focus: Governance, Sustainability, and “The Seed Bank” model.
  • Teaching Point: Mastodon represents the “Third Way” (Federation). It shifts the problem from “Corporate Benevolence” to “Community Maintenance.”
  • Discussion: Is Federation a viable solution for long-term preservation? (Pros: No single point of failure. Cons: No single point of funding).

VI. Assessment Strategies §

The Portfolio Model §

Move away from exams. Multiple-choice tests cannot measure preservation skills. Grade based on the Field Report—documentation of actual preservation work.

Rubric: Platform Autopsy §

  • Taxonomy Application (25%): Does the student correctly identify Vivibytes, Umbrabytes, etc.?
  • Analysis Depth (30%): Does the diagnosis of “Cause of Death” go beyond the press release? (e.g., analyzing the acqui-hire intent).
  • Historical Accuracy (20%): Is the timeline of the platform correct?
  • Recommendations (15%): Are the suggestions for future preservation actionable?
  • Writing/Clarity (10%): Is the report professional and accessible?

Apparatus • Reference, Syllabi & Curricular Toolkit

Appendix E: Professional Resources for Archaeobytologists

4 min read 687 words

Introduction §

Since “Archaeobytologist” is not yet a standard job title in most HR databases, this appendix functions as a “Translation Guide” for students entering the workforce. It maps the skills learned in the Archaeobytology curriculum to existing job market sectors, while also providing a roadmap for the future institutionalization of the field.


I. The Career Tracks (The “Translation” Layer) §

Source: Chapter 16

Track 1: Memory Institution Practitioner §

  • The Role: The custodian working within established libraries and archives.
  • Current Job Titles: Digital Archivist, Born-Digital Specialist, Metadata Librarian, Digital Curation Officer.
  • Target Employers: National archives (NARA, The National Archives UK), university libraries, special collections, museums (The Strong, Rhizome).
  • Key Skill Translation: Triage (Ch 5) -> “Appraisal”; Forensics (Ch 8) -> “Bit-level Preservation.”

Track 2: Industry Sovereignty Architect §

  • The Role: The builder working inside tech companies to enable data portability and ethical governance.
  • Current Job Titles: Data Portability Lead, Trust & Safety Policy Manager, Site Reliability Engineer (SRE), Open Source Program Office (OSPO) Manager.
  • Target Employers: Tech platforms (specifically Governance/Export teams), Mozilla, DuckDuckGo, Federated social networks (Ghost, Mastodon hosts).
  • Key Skill Translation: Sovereignty Design (Ch 12) -> “User Trust & Safety”; Distributed Governance (Ch 13) -> “Decentralization Strategy.”

Track 3: The Preservation Consultant (The “Anvil” Track) §

  • The Role: The mercenary expert helping organizations navigate digital death or transition.
  • Current Job Titles: Digital Asset Manager (DAM), Information Governance Consultant, Legacy System Migration Specialist.
  • Target Clients: Non-profits closing down, law firms (eDiscovery), legacy media companies digitizing back catalogues.
  • Key Skill Translation: Excavation (Ch 7) -> “Data Migration”; Triage (Ch 5) -> “Information Lifecycle Management.”

Track 4: The Public Advocate §

  • The Role: The activist fighting for the legal right to preserve.
  • Current Job Titles: Technology Policy Analyst, Digital Rights Campaigner, Campaign Director.
  • Target Employers: EFF, Fight for the Future, Creative Commons, Public Knowledge.
  • Key Skill Translation: Custodial Filter (Ch 9) -> “Digital Rights Policy”; Movement Building (Ch 16) -> “Advocacy.”

II. Professional Societies & “Trading Zones” §

Before the “Society for Archaeobytology” is formally established (Year 4 Goal), these are the spaces where the work currently happens.

  • The Maintainers: A global research network interested in the concepts of maintenance, infrastructure, and repair.
  • National Digital Stewardship Alliance (NDSA): A consortium of institutions committed to the long-term preservation of digital information.
  • iPres: The International Conference on Digital Preservation (the premier venue for technical preservation work).
  • Association of Internet Researchers (AoIR): For the cultural/social analysis of platforms.
  • Society for Social Studies of Science (4S): For the political economy and STS aspects of the curriculum.

III. Funding & Grant Sources §

For students taking the “Institution Building” track (Chapter 11), these are the primary capital sources for preservation work.

Public/Government:

  • NEH (Office of Digital Humanities): For cultural heritage projects.
  • IMLS (Institute of Museum and Library Services): For infrastructure and access.
  • NSF: For technical infrastructure (though often requires CS partnership).

Private Philanthropy:

  • Mellon Foundation: The largest funder of digital preservation and scholarly communications.
  • Sloan Foundation: Funds “Universal Access to Knowledge” projects.
  • Filecoin Foundation for the Decentralized Web: Funding for distributed/p2p preservation architectures.

IV. The “Certified Archaeobytologist” Roadmap §

Source: Chapter 16

This section outlines the proposed professional credential discussed in the movement-building chapter.

  • Vision: A formal credential validating expertise in Triage (rapid decision making), Forensics (technical recovery), and Sovereignty Design (ethical architecture).
  • Current Equivalent: Students seeking validation today should look to the Academy of Certified Archivists (ACA) combined with the SAA Digital Archives Specialist (DAS) certificate.

Resources for practitioners facing legal threats (DMCA) or ethical crises.

  • Electronic Frontier Foundation (EFF) Coders’ Rights Project: Legal assistance for researchers and archivists facing reverse-engineering or scraping threats.
  • The Copyright Office (US) Section 1201 Exemptions: The specific triennial rule-making process where archivists must fight for the right to break DRM for preservation.
  • Lawyers for Good Government: Pro bono legal support for public interest technology work.

Apparatus • Reference, Syllabi & Curricular Toolkit

Bibliography: Bibliography & Works Cited

14 min read 3,049 words

Core Archaeobytology Texts §

Foundational Theory §

Derrida, Jacques. Archive Fever: A Freudian Impression. University of Chicago Press, 1996.

  • Philosophical meditation on archives, memory, and the death drive. Essential for understanding archival impulse.

Ernst, Wolfgang. Digital Memory and the Archive. University of Minnesota Press, 2013.

  • Media archaeological perspective on digital preservation. Bridges theory and technical practice.

Kirschenbaum, Matthew G. Mechanisms: New Media and the Forensic Imagination. MIT Press, 2008.

  • Foundational text on digital forensics and materiality. Demonstrates how to study digital artifacts as physical objects.

Parikka, Jussi. What Is Media Archaeology? Polity, 2012.

  • Concise introduction to media archaeology. Shows how to excavate dead media theoretically.

Chun, Wendy Hui Kyong. “The Enduring Ephemeral, or the Future Is a Memory.” Critical Inquiry 35, no. 1 (2008): 148-171.

  • Theorizes the paradox of digital permanence/ephemerality. Essential for understanding digital mortality.

Platform Studies and Critique §

Gillespie, Tarleton. Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions That Shape Social Media. Yale University Press, 2018.

  • How platforms curate, moderate, and control content. Essential for understanding platform power.

Noble, Safiya Umoja. Algorithms of Oppression: How Search Engines Reinforce Racism. NYU Press, 2018.

  • Critical analysis of algorithmic bias. Shows why platform design is political.

Zuboff, Shoshana. The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power. PublicAffairs, 2019.

  • Comprehensive critique of surveillance business models. Explains why platforms murder culture for profit.

Doctorow, Cory. The Internet Con: How to Seize the Means of Computation. Verso, 2023.

  • Advocacy for interoperability and user sovereignty. Practical vision for alternatives to platform capitalism.

Rushkoff, Douglas. Team Human. W.W. Norton, 2019.

  • Humanistic critique of platform society. Argues for human-centered technology.

Digital Preservation and Archiving §

Brügger, Niels, and Ralph Schroeder, eds. The Web as History: Using Web Archives to Understand the Past and the Present. UCL Press, 2017.

  • Case studies of web preservation projects. Shows how to use archived web data for historical research.

Ankerson, Megan Sapnar. Dot-com Design: The Rise of a Usable, Social, Commercial Web. NYU Press, 2018.

  • History of early web design using archived sites. Demonstrates value of preserved digital culture.

Brügger, Niels. “Website History and the Website as an Object of Study.” New Media & Society 11, no. 1-2 (2009): 115-132.

  • Theorizes websites as historical objects. Methodological framework for studying archived sites.

Manoff, Marlene. “Theories of the Archive from Across the Disciplines.” Portal: Libraries and the Academy 4, no. 1 (2004): 9-25.

  • Survey of archival theory across disciplines. Shows how different fields understand archives.

Cook, Terry. “What is Past is Prologue: A History of Archival Ideas Since 1898, and the Future Paradigm Shift.” Archivaria 43 (1997): 17-63.

  • Evolution of archival theory. Essential for understanding contemporary preservation practices.

Digital Sovereignty and Infrastructure §

Foundational Texts §

Lessig, Lawrence. Code: Version 2.0. Basic Books, 2006.

  • “Code is law” - how digital architecture shapes behavior. Essential for understanding sovereignty.

Schneier, Bruce. Data and Goliath: The Hidden Battles to Collect Your Data and Control Your World. W.W. Norton, 2015.

  • Comprehensive analysis of surveillance and data collection. Practical guide to digital security.

Véliz, Carissa. Privacy Is Power: Why and How You Should Take Back Control of Your Data. Melville House, 2020.

  • Accessible argument for data sovereignty. Bridges philosophy and practice.

Schneier, Bruce. Click Here to Kill Everybody: Security and Survival in a Hyper-connected World. W.W. Norton, 2018.

  • Internet of Things security risks. Shows vulnerabilities in digital infrastructure.

Commons and Collective Governance §

Ostrom, Elinor. Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press, 1990.

  • Nobel Prize-winning framework for commons governance. Essential for understanding Seed Bank design.

Benkler, Yochai. The Wealth of Networks: How Social Production Transforms Markets and Freedom. Yale University Press, 2006.

  • Theory of peer production and commons-based alternatives to capitalism.

Bollier, David. Think Like a Commoner: A Short Introduction to the Life of the Commons. New Society Publishers, 2014.

  • Accessible introduction to commons thinking. Shows alternatives to private/state ownership.

Hess, Charlotte, and Elinor Ostrom, eds. Understanding Knowledge as a Commons: From Theory to Practice. MIT Press, 2006.

  • Applies commons theory to information and knowledge. Directly relevant to digital preservation.

Decentralization and Protocols §

Bauwens, Michel, and Vasilis Kostakis. Network Society and Future Scenarios for a Collaborative Economy. Palgrave Macmillan, 2014.

  • Theory of peer-to-peer networks and collaborative commons.

Baran, Paul. “On Distributed Communications Networks.” IEEE Transactions on Communications Systems 12, no. 1 (1964): 1-9.

  • Original distributed network design (ARPANET precursor). Historical foundation for decentralization.

Staltz, André. “The Web Began Dying in 2014, Here’s How.” Blog post, 2017. https://staltz.com/the-web-began-dying-in-2014-heres-how.html

  • Accessible critique of platform centralization. Documents shift from open to closed web.

Ethics and Social Justice §

Archival Ethics §

Caswell, Michelle. Urgent Archives: Enacting Liberatory Memory Work. Routledge, 2021.

  • Framework for ethical archiving centered on social justice and community needs.

Caswell, Michelle. “Seeing Yourself in History: Community Archives and the Fight Against Symbolic Annihilation.” The Public Historian 36, no. 4 (2014): 26-37.

  • How archives can counter erasure of marginalized communities.

Jimerson, Randall C. Archives Power: Memory, Accountability, and Social Justice. Society of American Archivists, 2009.

  • Comprehensive treatment of archives as instruments of power and justice.

Flinn, Andrew. “Community Histories, Community Archives: Some Opportunities and Challenges.” Journal of the Society of Archivists 28, no. 2 (2007): 151-176.

  • Community-led archiving as alternative to institutional control.

Nissenbaum, Helen. Privacy in Context: Technology, Policy, and the Integrity of Social Life. Stanford University Press, 2009.

  • “Contextual integrity” framework for privacy. Essential for understanding preservation ethics.

Solove, Daniel J. Nothing to Hide: The False Tradeoff Between Privacy and Security. Yale University Press, 2011.

  • Argues against “nothing to hide” argument. Shows why privacy matters even for ordinary people.

Rosen, Jeffrey. “The Right to Be Forgotten.” Stanford Law Review Online 64 (2012): 88-92.

  • Legal and ethical dimensions of deletion rights. Relevant to preservation consent issues.

Cohen, Julie E. “What Privacy Is For.” Harvard Law Review 126, no. 7 (2013): 1904-1933.

  • Theorizes privacy as essential for human flourishing, not just individual right.

Technology and Justice §

Benjamin, Ruha. Race After Technology: Abolitionist Tools for the New Jim Code. Polity, 2019.

  • How technology encodes racism. Essential for understanding bias in preservation decisions.

Costanza-Chock, Sasha. Design Justice: Community-Led Practices to Build the Worlds We Need. MIT Press, 2020.

  • Framework for justice-centered design. Applicable to building sovereign systems.

D’Ignazio, Catherine, and Lauren F. Klein. Data Feminism. MIT Press, 2020.

  • Feminist approach to data and technology. Shows how to center marginalized perspectives.

Discipline Formation and Movement Building §

How Disciplines Form §

Abbott, Andrew. Chaos of Disciplines. University of Chicago Press, 2001.

  • Sociological analysis of academic disciplines. Shows how fields compete and evolve.

Klein, Julie Thompson. Interdisciplining Digital Humanities: Boundary Work in an Emerging Field. University of Michigan Press, 2015.

  • Case study of Digital Humanities discipline formation. Direct parallel to Archaeobytology.

Kuhn, Thomas S. The Structure of Scientific Revolutions. University of Chicago Press, 1962 [1996].

  • Classic on paradigm shifts. Relevant to understanding how new disciplines emerge.

Small, Mario Luis. “How to Conduct a Mixed Methods Study: Recent Trends in a Rapidly Growing Literature.” Annual Review of Sociology 37 (2011): 57-86.

  • Methodological pluralism in emerging fields.

Boundary Work §

Gieryn, Thomas F. “Boundary-Work and the Demarcation of Science from Non-Science: Strains and Interests in Professional Ideologies of Scientists.” American Sociological Review 48, no. 6 (1983): 781-795.

  • How disciplines define themselves through exclusion. Essential for understanding disciplinary boundaries.

Star, Susan Leigh, and James R. Griesemer. “Institutional Ecology, ‘Translations’ and Boundary Objects: Amateurs and Professionals in Berkeley’s Museum of Vertebrate Zoology, 1907-39.” Social Studies of Science 19, no. 3 (1989): 387-420.

  • How interdisciplinary work creates “boundary objects.” Relevant to Archaeobytology’s synthetic nature.

Public Scholarship §

Burawoy, Michael. “For Public Sociology.” American Sociological Review 70, no. 1 (2005): 4-28.

  • Advocacy for scholarship engaging public, not just academy. Model for public Archaeobytology.

Posner, Miriam. “Here and There: Creating DH Community.” In Debates in the Digital Humanities 2016, edited by Matthew K. Gold and Lauren F. Klein, 265-276. University of Minnesota Press, 2016.

  • Building scholarly community in interdisciplinary field. Practical lessons for Archaeobytology.

Technical Methods and Tools §

Web Archiving and Scraping §

Brügger, Niels. Web Historiography and Internet Studies. Polity, 2018.

  • Methodological framework for studying archived web. Technical and theoretical synthesis.

Milligan, Ian. History in the Age of Abundance? How the Web Is Transforming Historical Research. McGill-Queen’s University Press, 2019.

  • Practical guide to using web archives for historical research. Shows tools and methods.

Archive Team Wiki. https://wiki.archiveteam.org/

  • Community-maintained documentation of preservation methods. Primary source and technical manual.

Digital Forensics §

Kirschenbaum, Matthew G., Richard Ovenden, and Gabriela Redwine. Digital Forensics and Born-Digital Content in Cultural Heritage Collections. Council on Library and Information Resources (CLIR), 2010.

  • Practical guide to digital forensics for archivists and historians. Groundbreaking report on applying forensic methods to archives.

Carrier, Brian. File System Forensic Analysis. Addison-Wesley, 2005.

  • Technical manual for file system forensics. Advanced but comprehensive.

Emulation and Preservation §

Guttenbrunner, Mark, Andreas Rauber, and Christoph Becker. “Evaluating Strategies for the Preservation of Console Video Games.” International Journal on Digital Libraries 11, no. 1 (2010): 37-60.

  • Technical strategies for emulation. Video game preservation as case study.

Rothenberg, Jeff. “Avoiding Technological Quicksand: Finding a Viable Technical Foundation for Digital Preservation.” Council on Library and Information Resources, 1999.

  • Classic argument for emulation over migration. Technical preservation strategy.

Political Economy and Critique §

Platform Capitalism §

Srnicek, Nick. Platform Capitalism. Polity, 2016.

  • Economic analysis of platform business models. Shows why platforms are structurally extractive.

Duffy, Brooke Erin. (Not) Getting Paid to Do What You Love: Gender, Social Media, and Aspirational Work. Yale University Press, 2017.

  • How platforms exploit creative labor. Shows human cost of platform capitalism.

Scholz, Trebor, ed. Digital Labor: The Internet as Playground and Factory. Routledge, 2012.

  • Collection on labor in digital platforms. Shows extraction mechanisms.

Alternative Economics §

Scholz, Trebor, and Nathan Schneider, eds. Ours to Hack and to Own: The Rise of Platform Cooperatives. OR Books, 2016.

  • Collection on cooperative alternatives to platform capitalism. Practical models for Anvil economics.

Schneider, Nathan. “An Internet of Ownership: Democratic Design for the Online Economy.” The Sociological Review 68, no. 2 (2020): 320-340.

  • Platform cooperatives as sovereignty model. Bridges theory and practice.

Bauwens, Michel. “The Political Economy of Peer Production.” Post-autistic Economics Review 37 (2006): 33-44.

  • Economic theory of peer production. Alternative to market and state.

Tech Policy and Regulation §

Wu, Tim. The Master Switch: The Rise and Fall of Information Empires. Knopf, 2010.

  • Historical cycles of open/closed information systems. Shows patterns in tech consolidation.

Pasquale, Frank. The Black Box Society: The Secret Algorithms That Control Money and Information. Harvard University Press, 2015.

  • Critique of algorithmic opacity. Argues for transparency and accountability.

Crawford, Kate. Atlas of AI: Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press, 2021.

  • Material and political economy of AI. Shows infrastructure behind “cloud” computing.

Craft, Making, and Building §

Philosophy of Making §

Sennett, Richard. The Craftsman. Yale University Press, 2008.

  • Philosophy of skilled practice and making. Relevant to Anvil as craft practice.

Pye, David. The Nature and Art of Workmanship. Cambridge University Press, 1968.

  • Classic text on craft and workmanship. Distinguishes workmanship of risk from workmanship of certainty.

Crawford, Matthew B. Shop Class as Soulcraft: An Inquiry into the Value of Work. Penguin, 2009.

  • Argument for hands-on work. Relevant to building vs. theorizing tension.

Critical Making §

Ratto, Matt. “Critical Making: Conceptual and Material Studies in Technology and Social Life.” The Information Society 27, no. 4 (2011): 252-260.

  • Framework for making as research and critique. Bridges scholarship and building.

Hertz, Garnet. Critical Making: Software Studies. 2012. http://www.conceptlab.com/criticalmaking/

  • Collection on making as intellectual practice. Shows how building produces knowledge.

Historical Context and Case Studies §

Internet History §

Abbate, Janet. Inventing the Internet. MIT Press, 1999.

  • Comprehensive history of internet development. Essential context for understanding current crisis.

Hafner, Katie, and Matthew Lyon. Where Wizards Stay Up Late: The Origins of the Internet. Simon & Schuster, 1996.

  • Accessible history of ARPANET. Shows original decentralized vision.

Turner, Fred. From Counterculture to Cyberculture: Stewart Brand, the Whole Earth Network, and the Rise of Digital Utopianism. University of Chicago Press, 2006.

  • How 1960s counterculture shaped internet ideology. Explains libertarian tech culture.

Platform Histories §

boyd, danah. It’s Complicated: The Social Lives of Networked Teens. Yale University Press, 2014.

  • Ethnography of teen social media use. Shows what’s at stake when platforms die.

Marwick, Alice E. Status Update: Celebrity, Publicity, and Branding in the Social Media Age. Yale University Press, 2013.

  • How social media transforms identity and labor. Shows platform effects on culture.

Baym, Nancy K. Personal Connections in the Digital Age. Polity, 2015 (2nd ed.).

  • How digital platforms shape relationships and community. Essential context.

Specific Platform Studies §

Salter, Anastasia, and Bridget Blodgett. Toxic Geek Masculinity in Media: Sexism, Trolling, and Identity Policing. Palgrave Macmillan, 2017.

  • Case study of platform culture (Reddit, 4chan, gaming). Shows dark side of platforms.

Bucher, Taina. If…Then: Algorithmic Power and Politics. Oxford University Press, 2018.

  • How algorithms shape platform experience. Essential for understanding platform design.

Relevant Adjacent Fields §

Science and Technology Studies (STS) §

Latour, Bruno. Reassembling the Social: An Introduction to Actor-Network-Theory. Oxford University Press, 2005.

  • Actor-network theory framework. Useful for analyzing sociotechnical systems.

Winner, Langdon. “Do Artifacts Have Politics?” Daedalus 109, no. 1 (1980): 121-136.

  • Classic essay on how technology embeds politics. Essential for understanding sovereignty architecture.

Bijker, Wiebe E., Thomas P. Hughes, and Trevor Pinch, eds. The Social Construction of Technological Systems. MIT Press, 1987.

  • Foundational collection in STS. Shows how technology and society co-constitute each other.

Library and Information Science §

Buckland, Michael. “Information as Thing.” Journal of the American Society for Information Science 42, no. 5 (1991): 351-360.

  • Theorizes information as physical/digital object. Relevant to artifact preservation.

Dourish, Paul. The Stuff of Bits: An Essay on the Materialities of Information. MIT Press, 2017.

  • Materiality of digital information. Bridges computer science and cultural studies.

Huvila, Isto. “Participatory Archive: Towards Decentralised Curation, Radical User Orientation, and Broader Contextualisation of Records Management.” Archival Science 8, no. 1 (2008): 15-36.

  • Community-led archiving models. Alternative to institutional control.

Media Studies §

Jenkins, Henry. Convergence Culture: Where Old and New Media Collide. NYU Press, 2006.

  • How media platforms shape participatory culture. Shows cultural stakes of platforms.

Papacharissi, Zizi. A Private Sphere: Democracy in a Digital Age. Polity, 2010.

  • How social media reshapes public/private boundaries. Relevant to preservation ethics.

van Dijck, José. The Culture of Connectivity: A Critical History of Social Media. Oxford University Press, 2013.

  • Critical history of social media platforms. Documents shift from user-centered to corporate-centered.

Primary Sources and Documentation §

Organizational Documents §

Internet Archive. “About the Internet Archive.” https://archive.org/about/

  • Mission statement and organizational structure of world’s largest digital archive.

Archive Team. “Who We Are.” https://archiveteam.org/

  • Documentation of guerrilla archiving practices and community.

Electronic Frontier Foundation. “About EFF.” https://www.eff.org/about

  • Leading digital rights organization. Source for policy advocacy models.

Technical Standards and Protocols §

W3C. “ActivityPub.” https://www.w3.org/TR/activitypub/

  • Federated social networking protocol specification. Technical foundation for distributed platforms.

IETF. “SMTP RFC 5321.” https://tools.ietf.org/html/rfc5321

  • Email protocol specification. Example of successful open protocol.

IPFS. “InterPlanetary File System Documentation.” https://docs.ipfs.tech/

  • Distributed file storage protocol. Alternative to centralized hosting.

Community Resources §

IndieWeb Wiki. https://indieweb.org/

  • Community documentation of sovereign web practices. Primary source for self-hosting methods.

Mastodon Documentation. https://docs.joinmastodon.org/

  • Federated social network documentation. Technical and community governance resources.

Flashpoint Archive Project. http://flashpointproject.github.io/

  • Flash game preservation project. Case study and technical resource.

Manifestos and Calls to Action §

Kahle, Brewster. “Preserving the Internet.” Scientific American 276, no. 3 (1997): 82-83.

  • Early call for comprehensive web preservation. Founding vision of Internet Archive.

Doctorow, Cory. “Adversarial Interoperability.” EFF, 2019. https://www.eff.org/deeplinks/2019/10/adversarial-interoperability

  • Argument for legal right to make platforms interoperate. Policy advocacy framework.

Çelik, Tantek. “Own Your Data.” https://indieweb.org/own_your_data

  • IndieWeb manifesto for data ownership. Accessible articulation of sovereignty principles.

Stallman, Richard. “The GNU Manifesto.” 1985. https://www.gnu.org/gnu/manifesto.html

  • Founding document of free software movement. Historical precedent for digital sovereignty.

Barlow, John Perry. “A Declaration of the Independence of Cyberspace.” Electronic Frontier Foundation, 1996. https://www.eff.org/cyberspace-independence

  • Utopian vision of internet freedom. Historical document showing early sovereignty thinking (critiquable but influential).

Further Resources §

Podcasts and Media §

Your Undivided Attention (Center for Humane Technology)

  • Critical analysis of platform design and addiction. Accessible to general audiences.

Recode Media (Vox)

  • Tech journalism covering platform politics. Current events and industry analysis.

The Download (MIT Technology Review)

  • Daily tech news with critical perspective.

Blogs and Online Writing §

Cory Doctorow’s Pluralistic - https://pluralistic.net/

  • Daily blog on tech policy, platforms, and digital rights. Essential reading.

Anil Dash’s Blog - https://anildash.com/

  • Tech industry insider with critical perspective on platforms.

Darius Kazemi’s Blog - https://tinysubversions.com/

  • Developer building alternative platforms and bots. Practical sovereignty projects.

Video Resources §

Brewster Kahle TEDx Talks

  • Internet Archive founder on preservation mission. Accessible introductions.

Documentaries:

  • Downloaded (2013) - Napster history, shows platform life cycle
  • The Cleaners (2018) - Content moderation labor, shows platform power
  • The Social Dilemma (2020) - Platform critique (populist but accessible)

Conclusion §

This bibliography represents the intellectual foundations of Archaeobytology—drawing from archives, computer science, philosophy, political economy, craft, law, and activism. No single discipline provides all the tools needed; Archaeobytology synthesizes them.

Recommended Starting Points:

For theory: Derrida, Chun, Kirschenbaum, Parikka For practice: Archive Team Wiki, Brügger, Milligan For politics: Doctorow, Zuboff, Lessig, Schneier For ethics: Caswell, Nissenbaum, Benjamin For building: Benkler, Ostrom, Schneider, Sennett

Next Steps:

  1. Read broadly across disciplines (don’t stay in one silo)
  2. Follow practitioners on social media (Twitter, Mastodon, blogs)
  3. Join communities (Archive Team, IndieWeb, federated platforms)
  4. Build something (tools, archives, protocols)
  5. Teach others (write, speak, organize)

Archaeobytology is a young discipline. This bibliography will grow as the field develops. Add to it. Challenge it. Build on it.

Now go preserve something.


End of Bibliography

Apparatus • Reference, Syllabi & Curricular Toolkit

Index: Subject & Concept Index

9 min read 1,923 words

Core Concepts §

Archaeobyte - Digital artifact that was once alive, died through platform shutdown or obsolescence, and has been preserved in some form

Archaeobytology - The study and practice of excavating, preserving, interpreting, and building with digital artifacts, particularly those murdered by platforms

Anvil, The - The creative/building practice of Archaeobytology; forging tools, protocols, and institutions that embody digital sovereignty

Archive, The - The preservation/memory practice of Archaeobytology; excavating and maintaining murdered digital artifacts

Bit Rot - Gradual degradation of digital storage media leading to data loss

Chain of Custody - Documentation of who handled an artifact and when, essential for forensic integrity

Context Collapse - When content created for specific audience becomes accessible to unintended audiences (e.g., private forum posts made public in archive)

Custodial Filter, The - Five-question ethical framework for triage decisions (significance, fragility, feasibility, redundancy, ethics)

Digital Ground - Infrastructure and storage a user controls; Third Pillar of sovereignty

Digital Sovereignty - Ability to exist, communicate, and build in digital space without corporate gatekeeping; embodied in Three Pillars

Dual Soul - Archaeobytology’s integrated practice of preservation (Archive) and creation (Anvil)

Emulation - Running old software/platforms on modern systems by simulating original hardware/OS

Format Migration - Converting files from obsolete formats to current standards

Link Rot - URLs breaking over time as sites move, reorganize, or disappear

Murdered Platform - Platform deliberately killed by corporate decision, not natural obsolescence

Petribyte - Digital artifact so old and well-preserved it’s achieved monument status (like stone tablets)

Platform Death - Shutdown of digital platform resulting in loss of hosted content and communities

POSSE - “Post On your Site, Syndicate Elsewhere” - IndieWeb practice of owning original content

Provenance - Documentation of artifact’s origin, creation, and history

Shadow Preservation - Archiving content without explicit permission, often in legal gray areas

Stratigraphic Analysis - Studying layers of digital artifacts to understand temporal and contextual relationships

Surveillance Capitalism - Business model extracting behavioral data for profit (Zuboff)

Triage - Systematic methodology for deciding what to preserve when resources are scarce

Umbrabyte - Digital artifact that’s technically dead but exists in fragmentary form; haunting but not fully preserved

Vivibyte - Digital artifact currently alive but endangered by platform instability

Web Scraping - Automated extraction of website content for preservation


Three Pillars of Digital Sovereignty §

Pillar 1: Declaration (I Am) - Self-owned identity and persistent presence without platform permission

Pillar 2: Connection (Instant Message) - Direct communication and portable relationships without corporate intermediation

Pillar 3: Ground (Digital Real Estate) - Owned infrastructure, data, and domains; ability to migrate without loss


Four Institutions §

The Archive - Preservation organization; saves murdered platforms and maintains artifacts for decades

The Anvil - Profitable business building sovereign tools/platforms while embodying Three Pillars

The Seed Bank - Distributed commons governance structure; peer-to-peer preservation without single point of failure

The Haunted Forest - Memory institution (museum/memorial); curates and interprets preserved artifacts for public


Key Organizations §

Archive Team - Guerrilla digital archiving collective that mobilizes to rescue dying platforms

Creative Commons - Organization providing open licensing frameworks for content sharing

Electronic Frontier Foundation (EFF) - Digital rights advocacy organization

Internet Archive - Non-profit digital library providing free access to websites, books, media, and software

Library of Congress - U.S. national library with extensive digital preservation programs

Mozilla Foundation - Non-profit supporting open web technologies and user sovereignty

Wikimedia Foundation - Operates Wikipedia and sister projects; model of commons governance


Technologies and Protocols §

ActivityPub - W3C standard for federated social networking (used by Mastodon, Pixelfed, PeerTube)

Archive-It - Web archiving service provided by Internet Archive

ArchiveBox - Open-source self-hosted web archiving tool

BitTorrent - Peer-to-peer file sharing protocol used for distributed preservation

DNS (Domain Name System) - Internet addressing system (centralized, vulnerable to control)

Emularity - JavaScript emulator framework allowing old software to run in web browsers

ENS (Ethereum Name Service) - Blockchain-based naming system for decentralized identity

Flashpoint - Preservation project saving Flash games and animations

Ghost - Open-source publishing platform supporting custom domains and data export

IPFS (InterPlanetary File System) - Peer-to-peer distributed file system for permanent web storage

Matrix - Open protocol for federated, end-to-end encrypted communication

Mastodon - Federated social network using ActivityPub protocol

Nextcloud - Open-source self-hosted productivity platform (alternative to Google Workspace)

Obsidian - Knowledge management app storing files locally in Markdown (data sovereignty)

RSS (Really Simple Syndication) - Open protocol for content syndication and subscriptions

Signal - End-to-end encrypted messaging app using Signal Protocol

WARC (Web ARChive format) - ISO standard file format for web archiving

Wayback Machine - Internet Archive’s web page archiving service (800+ billion pages)

WebRecorder - Tool for high-fidelity web archiving including dynamic content

WordPress - Open-source content management system powering 40%+ of web

wget - Command-line tool for downloading websites


Platforms (Murdered or Endangered) §

AOL (America Online) - Early internet service provider; email service declined/abandoned

Blogger - Google-owned blogging platform; free but corporate-controlled

Discord - Proprietary chat platform; communities at risk if platform shuts down

Ello - Anti-advertising social network; failed to achieve sustainability

Facebook/Meta - Social media monopoly; extractive business model, surveillance capitalism

Flickr - Photo sharing platform; multiple ownership changes threatened survival

FriendFeed - Social aggregation platform; acquired and killed by Facebook (2009)

GeoCities - Early web hosting service; murdered by Yahoo in 2009 (30 million sites lost)

Google+ - Google’s social network; shut down 2019

Google Reader - RSS feed reader; killed by Google 2013 despite millions of users

Instagram - Photo sharing owned by Meta; algorithmic feed, no data portability

LiveJournal - Blogging/social platform; Russian ownership drove user exodus

Medium - Publishing platform; multiple business model pivots, corporate control

Mixer - Game streaming platform; Microsoft shut down 2020

MySpace - Early social network; lost 12 years of music in 2019 server migration

Snapchat - Ephemeral messaging app; content designed to disappear

Substack - Newsletter platform; writers don’t own domains or full subscriber relationships

TikTok - Video sharing platform; facing potential bans, Chinese ownership controversy

Tumblr - Blogging platform; 2018 NSFW purge deleted millions of posts

Twitter/X - Microblogging platform; chaotic ownership under Musk, mass exodus

Vine - 6-second video platform; Twitter shut down 2017 (200 million videos at risk)

WhatsApp - Encrypted messaging owned by Meta; metadata surveillance, closed platform


Key Thinkers and Practitioners §

Benkler, Yochai - Scholar of peer production and commons-based alternatives

Bowker, Geoffrey C. - Information studies scholar; classification and infrastructure

Brewster Kahle - Founder of Internet Archive; digital preservation advocate

Caswell, Michelle - Archival studies scholar; community archives and social justice

Chun, Wendy Hui Kyong - Media studies scholar; digital memory and ephemerality

Doctorow, Cory - Science fiction author and digital rights activist; adversarial interoperability

Eugen Rochko - Creator of Mastodon federated social network

Gillespie, Tarleton - Media scholar studying platforms and content moderation

Kirschenbaum, Matthew - Digital humanities scholar; forensic approaches to digital artifacts

Lessig, Lawrence - Legal scholar; “code is law,” Creative Commons founder

Nissenbaum, Helen - Privacy scholar; contextual integrity framework

Noble, Safiya Umoja - Scholar of algorithmic bias and racism in technology

Ostrom, Elinor - Nobel laureate; commons governance frameworks

Parikka, Jussi - Media archaeology scholar; dead media studies

Schneier, Bruce - Security expert and cryptographer; surveillance and privacy

Star, Susan Leigh - Sociologist of science and infrastructure

Zuboff, Shoshana - Scholar of surveillance capitalism


DMCA (Digital Millennium Copyright Act) - U.S. law criminalizing circumvention of DRM; complicates preservation

Fair Use - Legal doctrine allowing limited use of copyrighted material without permission

GDPR (General Data Protection Regulation) - EU privacy law requiring data portability and deletion rights

Interoperability - Ability of different systems to communicate; essential for sovereignty

Platform Liability - Legal responsibility of platforms for user-generated content

Right to Archive - Proposed legal right to preserve digital content for historical purposes

Right to Be Forgotten - Legal right to request deletion of personal data (conflicts with preservation)

Right to Repair - Legal right to fix devices without manufacturer permission; relevant to digital sovereignty

Section 230 - U.S. law protecting platforms from liability for user content

Terms of Service (ToS) - Legal agreement users accept when joining platform; often restricts data ownership


Methodological Terms §

API Harvesting - Using platform APIs to bulk-download data for preservation

Chain-of-Custody Documentation - Recording who handled artifact and when; forensic integrity

Checksum/Hash - Mathematical signature verifying file hasn’t been altered

Crawler - Automated program systematically browsing and indexing web content

Defederation - Severing connections between federated instances due to moderation conflicts

Digital Forensics - Investigating digital artifacts to determine authenticity, provenance, and history

Metadata - Data about data (creation date, author, file type, etc.)

Robots.txt - File telling web crawlers which parts of site not to archive

Screen Scraping - Extracting visible content from websites (distinct from API access)

Site Mirroring - Creating complete local copy of website

Stratigraphic Excavation - Methodical layer-by-layer preservation documenting relationships between artifacts


Economic and Governance Models §

Cooperative (Co-op) - Business owned and controlled democratically by members/workers

Freemium - Free basic service with paid premium features

Open Core - Open-source base product with proprietary enterprise features

Platform Cooperative - Platform owned by users/workers rather than investors

Public Funding - Government grants, contracts, or direct funding for preservation

Subscription Model - Recurring payments for ongoing service/access

Venture Capital (VC) - Investment funding requiring exponential growth and exit (problematic for sovereignty)


Cultural and Community Terms §

Digital Humanities - Interdisciplinary field applying computational methods to humanities research

Fandom - Fan communities creating derivative works and cultural artifacts

IndieWeb - Movement promoting personal websites and data ownership

Media Archaeology - Field studying dead, obsolete, and imaginary media

Platform Studies - Examining how platform architectures shape culture and behavior

Science and Technology Studies (STS) - Interdisciplinary field studying science/tech and society

Slash Fiction - Fan fiction featuring romantic/sexual relationships; often LGBTQ+

Web 1.0 - Early web era (1990s-2000s) characterized by static pages and personal sites

Web 2.0 - Social web era (2000s-2010s) dominated by user-generated content on platforms

Web 3.0 - Contested term; blockchain advocates claim decentralized future; critics see financialization


Timeline of Platform Deaths §

1996 - ARPANET decommissioned (replaced by modern Internet)

2001 - Napster shut down by court order

2009 - GeoCities shut down by Yahoo (October 26)

2013 - Google Reader shut down (July 1)

2017 - Vine shut down by Twitter (January 17)

2018 - Tumblr NSFW purge (December 17)

2019 - Google+ shut down (April 2)

2019 - MySpace loses 12 years of music in server migration

2020 - Mixer shut down by Microsoft (July 22)

2022 - Twitter chaos begins under Musk ownership (October 27)


Core Questions §

“What should be preserved?” - Central triage question requiring ethical deliberation

“Who owns your identity?” - Question revealing platform control vs. user sovereignty

“Can you take your data with you?” - Test of true data ownership and portability

“What happens when the platform dies?” - Question exposing infrastructure vulnerability

“Who decides what the future can know about the past?” - Question of archival power and responsibility


Appendices Note §

For detailed tool instructions, see Appendix B: Tools & Resources For sample curricula, see Appendix C: Sample Syllabi For teaching materials, see Appendix D: Teaching Resources For career pathways, see Appendix E: Professional Resources For citations, see Bibliography


End of Index

This index provides quick reference to key concepts, organizations, technologies, and terms throughout the textbook. For definitions and context, consult the Glossary (Appendix A) or the chapters where terms first appear.