Unknown unknown #2 · Desktop app

Data Custodian, Not Data Owner

Unknown unknown #2 — the fences decide what your AI assistant can reach on its own, but they don’t tell you what you are permitted to put inside them. That’s a question about what you owe to other people, and you need to consider it carefully.

Using the Terminal (CLI) instead? Read the Terminal version →

The last post was about getting the fences in place — a sandbox for the commands Claude Code runs, and permission rules for the files it can open directly. Together they settle what the tool can get to on its own. But the fences leave untouched a prior question, the one no setting can answer for you: of the things you could place in front of Claude Code, which are actually yours to hand over?

For anyone who holds other people’s data — and most of us in universities, clinics, agencies, and firms do — this is a governance question. As a university-based researcher and teacher, much of the material on my computer is not, in the sense that matters here, mine. It was entrusted to me: by students, by research participants, by the people whose lives my work is about. What I owe them does not disappear just because an AI assistant has made it frictionless to share their information.

This obligation binds everyone who touches an AI assistant. The colleague who only ever opens Claude in a browser still has the same duty to the people whose data they hold in their files. And what you owe doesn’t change depending on which tool you use, be it Claude, ChatGPT, Gemini, or any other AI assistant.

Consent was given to you, not to the machine

The data you hold about other people came to you under terms. Your students’ records, your subjects’ interviews, your patients’ histories, your clients’ files — the people they describe agreed to your handling of them, for your stated purposes, within a relationship they understood. They did not agree to a third-party artificial-intelligence company receiving that same material and processing it on its own servers. That was not on the table when they consented.

Entering their data into Claude Code is therefore not a neutral act of using a tool. It is a new disclosure — a handing-over to an entity your students and subjects never met and never authorized. When you open a file using Claude Code, it feels identical to opening the file yourself: that’s the friction being smoothed away. You’ve just given Claude Code’s parent company access to that data. Whether or not the data is later used to train a model (a matter for a later post), the disclosure itself has already happened the moment the file is accessed by Claude Code.

Your obligations are specific, and they predate your use of Claude Code

Using Claude Code to give feedback on student papers means a student’s name, ideas, and perhaps struggles are exposed to a company that was never party to any agreement with that student. Asking Claude Code to help draft a recommendation letter means someone’s career trajectory and your private assessment of their limitations go in, too. An interview transcript where a participant told you something sensitive was shared with you, not with an AI company’s servers. The common thread is that another person trusted you with their information for a specific purpose, and that purpose did not include this.

There are a number of cases where there are significant legal boundaries: FERPA covers student records, HIPAA requires a formal Business Associate Agreement before a vendor may handle patient information at all, IRB protocols govern research-participant data, and confidentiality obligations attach to client files, NDAs, and your institution’s own policies. The details differ by institution and jurisdiction, and I am not offering legal advice. But you are expected to know which of these rules apply before the relevant data goes near an AI assistant.

A rule you can use before you paste: classify first

A good place to start setting yourself some discipline on this issue is to sort your data before it goes near Claude Code, not after. A simple three-color rule does most of the work:

  • Green — yours, and safe to share. Your own drafts, your analysis, public documents, already-published material. These can go into Claude Code freely.
  • Yellow — sensitive, but yours. Unpublished work, internal documents, anything you would not want public but that you are entitled to decide about. Yellow calls for a brief check before you proceed: Are you using a Claude Code account with appropriate data-retention settings? Do the terms of that account cover this kind of content? And do you actually need to share this, or can you work around it?
  • Red — other people’s protected data. Identifiable student, subject, patient, or client information; anything under a confidentiality obligation. This does not go in without explicit authorization and, often, a formal institutional agreement.

One caution on the yellow category: replacing names with code numbers does not make data yellow. Swapping identifiers for codes is pseudonymization, not de-identification — the original identities are still recoverable from the key you hold. A research transcript where “Sarah Chen” has become “Participant 7” is not de-identified; it is still red if the underlying consent and protocol do not cover processing the data with Claude Code. The question is not whether you have swapped the names for a numeric code, but whether the people in the data authorized you for this particular use. That is a harder question than it looks, and for most research data the answer is “no” until formal steps are taken.

So: classify your data before it touches Claude Code. The color you assign decides what happens next, and the assignment is a judgment only you can make.

When the answer is red

A red classification is not the end of the road; it’s a redirection. The right next step is asking whether your institution has a vetted, contracted path — an approved tool, a signed data-processing agreement, or an amendment to your IRB protocol. Be prepared for the honest answer that no such path yet exists: most institutions have sorted this out for large platforms but not for specialized research software, and IRBs were designed to protect participants from researchers, not from researchers’ software stacks. Asking is still the right move — it surfaces the gap, and that email is both your minimum due diligence and how institutions begin to build the capacity to answer.

The takeaway

The lines for the checklist beside your keyboard:

Before any file goes near Claude Code, classify it — green, yellow, or red — and ask two questions of anything that isn’t plainly yours: 1. Is this mine to share with a third party? 2. Did the people this data is about agree to this?

If the honest answer to either is “no” or “I’m not sure,” don’t let Claude Code touch it until you have asked someone whose job it is to know.

The cost of this one. Unlike the fence, there’s no clear setup for governance: it’s a habit of judgment, exercised before the data moves rather than after. For genuinely protected data, the task is one email to your IRB, compliance office, or institutional counsel, asking whether an approved path exists. What the answer reveals — that a vetted path exists, that one is being built, or that no one has asked yet — is exactly the information that can help protect both you and the people in your files.

Next: what actually leaves your computer when Claude Code reads a file — because even data you are fully entitled to use still travels somewhere, and where it goes is the next thing worth knowing.

Next in the series
Where Your Data Goes While Using Claude Code
Unknown unknown #3 · Desktop app version →