Skip to content

Build a Lifelog That Still Opens in 20 Years

A resilient lifelog is a personal data preservation system engineered to remain accessible and legible entirely independent of the proprietary software, cloud service, or hardware that produced it. That definition is deliberately strict. It rules out the sleep app with the beautiful charts, the account dashboard that renders a decade of steps in one scroll, and the vendor sync service that currently holds the only complete copy of a location history. Those are viewing tools. An archive is something else.

What Makes a Lifelog Survive Its Own Software

The structural weakness in most quantified-self practice is not the tracking hardware. Sensors fail predictably and get replaced. The failure that erases a decade of data is commercial: a service shuts down, pivots to enterprise customers, deprecates its v1 export endpoint, or quietly drops support for the file format it invented just about three product cycles ago. Each of those events is individually unremarkable. Across nearly twenty years, at least one of them is close to certain for any given vendor.

The useful design assumption, therefore, is pessimistic: every tracking application currently installed on a phone will be gone within approximately ten years. Adopting that premise changes the objective from keeping an account alive to keeping records independently readable. It also reorders the work. Instead of optimizing which app produces the nicest weekly summary, the practitioner inventories each source by what it actually records—discrete measurements, journal text, or location samples, and establishes how to export that source before another year of data accumulates behind it.

Saga and comparable lifelogging applications produce enormous volumes of passive context. The volume is the point. It is also what makes retroactive rescue expensive, because an inventory postponed for in the range of three years is an inventory of three unfamiliar schemas.

What Makes a Lifelog Survive Its Own Software

Export Routines and the Three Formats Worth Keeping

Scheduling comes before conversion. Set recurring exports for every source that matters, then translate each provider's files into plain text, CSV, or JSON. Logs that change frequently warrant an export every 7–14 days; slower-moving records tolerate longer gaps. Each run should be entered in an export log recording the export date, the source, and the date range covered, because the gap you cannot see is the gap you will discover in 2041.

Three formats clear the twenty-year bar: .txt,.csv, and.json. Their advantage is the absence of a dependency—no proprietary reader, no license server, no renderer. Archival institutions sort formats along roughly this axis of software independence; the Library of Congress digital preservation guidelines read that way.

Keep the original vendor export beside the converted copy. The pairing is what lets a future parser be checked against its source rather than trusted blindly. A conversion script should map source field names onto documented archive fields, not onto whatever the application happens to display this quarter; display labels change, field semantics change more slowly.

Converter Sanity Check

Test any new converter on a 1–3 day sample that deliberately contains a missing value, a free-text journal entry, and a measurement with units attached. If the script survives those three cases, run it across the full export. If it silently drops the null, it will silently drop thousands.

Worth stating plainly: this conversion preserves the raw timeline and discards the product. The polished screens, the animated graphs, the clever weekly comparisons—none of that comes along. What remains is a sortable sequence of observations. For a twenty-year horizon, that trade is the correct one, but it should be made with open eyes rather than discovered later as a disappointment.

One Sortable Time Field: UTC, ISO 8601, and the Offset Problem

Inconsistent time formatting destroys chronological integrity faster than any disk failure. Merge four sources that each write dates their own way—US month-first, ISO, Unix epoch, a localized string with an abbreviated timezone name, and the resulting archive sorts into nonsense. The damage is subtle because the files still open.

Choose one sortable time field before combining anything. Write each entry's instant in UTC as ISO 8601, in the form YYYY-MM-DDThh:mm:ssZ. Store the observed local offset in a separate field. Substituting local clock time for UTC is the single most common irreversible mistake in personal archives, because local time is ambiguous by design.

The ambiguity is concrete. During a daylight-saving transition, the local times 01:30 at −04:00 and 01:30 at −05:00 denote distinct instants: 2025-11-02T05:30:00Z and 2025-11-02T06:30:00Z. Stored as local clock time, those two readings sort identically, and an hour of heart-rate samples, journal entries, and location points can reverse order relative to one another.

A workable field set looks like timestamp_utc, local_offset, source, record_type, value, and unit, with the original source record identifier retained for duplicate detection across overlapping exports. One further annotation repays itself repeatedly: document whether a given source's timestamp marks the start of a measurement, its end, or the moment the application received it. Three vendors will use three conventions and none will tell you which.

Silent Corruption, SHA-256 Manifests, and the Key You Cannot Lose

Storage media degrade without announcing it. A bit flips on a long-idle drive, the file still opens, and the corruption surfaces years later as an impossible reading or a truncated row. Detection requires a reference computed while the data was known good.

The procedure is mechanical. Once an export is finalized, generate a SHA-256 checksum manifest covering its files. Hash before copying, then verify the destination against the manifest after the transfer completes. Keep a copy of the manifest apart from the backup it describes, since a manifest stored only inside a corrupted archive certifies nothing. Repeat verification every 6–12 months, and always after moving the archive to new storage. A mismatch names the specific degraded file, which is then restored from another verified copy.

Encryption belongs on off-site copies. A combined health, location, and journal timeline is among the most sensitive datasets an individual will ever hold, and off-site storage multiplies the parties who could touch it.

The Key Outlives Nothing

Encryption introduces a catastrophic single point of failure. A decryption key lost across twenty years renders every otherwise intact backup permanently unreadable. Keep decryption instructions and a recoverable key in a separate protected location, and describe the key-recovery procedure inside the archive documentation—without placing the key itself in the archive.

The boundary case deserves naming: a perfectly readable CSV sitting beside an encrypted off-site copy is not a twenty-year archive when the only decryption key lived on the phone that created it. None of these controls protects against sustained neglect, either; a manifest nobody recalculates is documentation of a past state, not evidence of a present one.

Silent Corruption, SHA-256 Manifests, and the Key You Cannot Lose

Three Copies, Rotating Media, and a README That Explains Itself

The classic 3-2-1 pattern adapts cleanly to personal archives: a working archive, a backup on separate local media, and an encrypted off-site copy. The geographic separation does the real work. Three drives sitting beside the same computer address one failure mode and ignore theft, fire, and flood.

Media rotation is the part most people skip. Review drive health and available capacity every 6–12 months, and migrate copies onto replacement hard drives or solid-state media on a planned 3–5 year cycle, verifying hashes after each move. The objective is to outpace hardware mortality rather than to discover it.

Then there is the README. A plain-text README.txt at the root of every archive copy is what allows a future reader—including the same person, two decades older and without the original devices or accounts, to interpret the files. It should document:

  • the directory layout and what lives where;
  • CSV columns or JSON fields, with units stated explicitly;
  • the UTC and local-offset convention, including per-source timestamp semantics;
  • source names, the export cadence used, and the conversion commands that produced the archive files;
  • known gaps, outages, and periods where a tracker was simply not worn.

Keep the README and the conversion scripts as plain text inside each copy. Scripts are documentation even when they no longer run; they record the transformation that was applied.

The Annual Restore Drill

Trace a single heart-rate reading end to end and the whole architecture becomes legible: it leaves the vendor export, gets mapped into documented fields, is normalized to timestamp_utc with its offset preserved, enters a hashed manifest, and must finally reappear—readable and correctly ordered, on a machine that has never seen the original app. That last step is the only one that proves the preceding four.

So the recommendation is unambiguous: a lifelog backup does not exist until it has been restored from scratch, and that restoration should be scheduled every twelve months without exception. Retrieve the off-site archive onto a transfer device. Move it to a new machine kept disconnected from any network. Using nothing but the raw files, the README, and recovered credentials, verify the hashes, decrypt, and rebuild a chronological timeline by hand.

Sample the oldest 30 days and the newest 30 days, plus a window containing a timezone-offset change. Confirm that at least one journal entry, one measurement with its unit, and one location sample can be read and sorted together by timestamp_utc. Log the drill date, which archive copy was tested, the checksum result, the decryption result, and every field that required knowledge not written down anywhere. Each of those undocumented fields is a defect; fix it in the README that week, while the memory still exists to fix it with.

Run the drill this year on a weekend you would otherwise lose to nothing in particular. An archive that has survived one honest restore is worth more than five years of exports that have never been opened.

Join Our Newsletter

Be the first to know.

No spam, just thoughtful updates.

Join the Conversation

No comments.

Write a Comment

Cookie settings