We Asked Ten Companies for Our Data

Ten services, ten access requests, ten archives of files no human was meant to read. The most interesting part of your data is the part that never arrives.

The email arrived on a Tuesday, the better part of a month after the request, in the flat, friendly register of automated correspondence: your data is ready. The link, it added, would expire soon. We clicked anyway, because this was the moment — the small séance at the center of digital life, when you finally come face to face with your data double. What descended was a ZIP file, which swelled into folders, then folders inside folders, named in the house style of an org chart we don’t work for. Whatever else was in there, the first thing it communicated was that nobody had arranged any of it with a person in mind.

Over the past few months, the desk filed access requests with ten services, invoking Article 15 of the General Data Protection Regulation — the EU law that, since 2018, has given anyone in Europe the right to a full view of what a company holds on them: the categories of data, the purposes, the recipients, the sources. This is not a scorecard. We won’t name the companies or rank their response times, because what turned out to be interesting wasn’t who performed but what the performance looks like. Think of this as fieldwork on an artifact: the data export, a document hundreds of millions of people are entitled to and almost nobody reads. The first thing to know about the genre is that its most interesting character never appears on stage. The files record, in exhausting detail, what you did. They say almost nothing about what was concluded from it.

Machine-readable, in the sense that a machine wrote it

What came back was consistent in spirit: archives of JSON and CSV, the occasional PDF of tables, timestamps in milliseconds since the first of January 1970, location stored as very large integers. When The Verge ran a version of this experiment in 2019, Google’s location history arrived as a single 61-megabyte JSON file with fields like “timestampMs” and “latitudeE7” — your latitude multiplied by ten million, so no machine ever has to cope with a decimal point. This is data that is machine-readable in the sense that a machine wrote it; another machine, somewhere, is the imagined audience.

Whether a person may read it is a separate question, and a murkier one than you’d expect. Article 12 of the GDPR promises information in a “concise, transparent, intelligible and easily accessible form, using clear and plain language,” but legal scholars note that this duty plausibly governs the communication around your request rather than the export itself. The explicit guarantee of a “structured, commonly used and machine-readable format” belongs to a different right entirely — Article 20’s data portability. A perfectly compliant export can be illegible to the person who asked for it. In a 2023 study titled “Needle in the Haystack,” researchers who exercised access rights across a set of services found roughly a quarter of the exports machine-readable, fewer than one in five understandable, and about one in nine both.

The long month

The statute gives companies a month — it says “one month,” not thirty days, a distinction that matters mostly to calendars. Some of our requests came back almost immediately, plainly the work of a script that had been waiting all along for someone to ask. Others used nearly the entire allotment. A month turns out to be a strange unit for this transaction: long enough to sever the request from whatever prompted it. You file in a flare of unease — an ad that knew too much, a recommendation that felt like being watched — and by the time the archive lands, the unease has metabolized into ordinary resignation. A 2024 USENIX Security study that filed access requests at five major platforms clocked response times from hours to weeks, and one researcher’s download link expired before they retrieved it, forcing a repeat request. The deadline functions less as consumer protection than as a cooling-off period for the impulse to ask.

The deeper absence is structural. What the archives contain, mostly, is testimony: every search, every tap, every ride, every pause. What they almost never contain is verdict — the inference layer, the categories the company has decided you belong to, the working assumptions about your income, your interests, your susceptibility. In a 2021 study, participants who examined their own exports from services including Spotify and YouTube found no explanation of how recommendations were determined — precisely the thing they most wanted — and the researchers noted that companies can argue inferred data contains trade secrets, leaving the required scope of access genuinely ambiguous. A 2025 follow-up comparing the same researchers’ 2018 and 2023 exports found that some companies had omitted data in earlier responses that surfaced in later ones, and that vague labelling made mapping the exports against the companies’ own privacy policies effectively impossible. The export, in other words, is not a window into the machine. It is a receipt printed by the machine, itemizing what it cost you to be processed.

There is a version of this story in which the fix is technical: better schemas, friendlier interfaces, the README-type file the 2021 researchers propose, describing each file and category so a person can at least navigate their own dossier. It would help, the way a map legend helps. But the honest conclusion of our ten requests is less comfortable. The right of access confirms that data exists; it does not give you leverage over it. You cannot edit the archive, appeal the inferences that aren’t in it, or slow the collection that continues while you read. Confirmation is not control, and a copy is not a say. On the Tuesday the last archive arrived, we opened the folders one more time and read our own clicks back to ourselves, timestamped to the millisecond — a perfect record of the wrong thing.