You are probably familiar with the term “unknown unknowns”—it is one category of a familiar epistemological taxonomy whose usage stretches from 19th century British references to a vaguely identified Persian proverb, to mid-twentieth century engineering, to former U.S. Secretary of Defense Donald Rumsfeld’s commentaries on intelligence gathering. A “known known” is settled. A “known unknown” is a gap you perceive, and therefore can act on: by asking, waiting, or investigating. But it is the unknown unknowns that do the damage—because they never announce themselves.
I’m a sociologist and a Professor at the Crown Family School of Social Work, Policy, and Practice at the University of Chicago. My research is primarily ethnographic: sustained participant-observation fieldwork and long interviews, aimed at learning how people in some particular world understand that world. Yet I have sometimes been motivated to answer sociological questions for which these methods are poorly suited. Rather than set these questions aside, I have collaborated with colleagues who are experts in other methods: regression analysis, multilevel modeling, network analysis, machine learning. When artificial intelligence tools like Claude Code became available, I wondered if they could help me pursue research using methods outside my personal toolbox. So I started to experiment.
Over the last three months, I’ve built some projects with Claude Code, talked about Claude Code with colleagues with substantial computing expertise, and conversed with Claude Code itself about what I’ve been doing. In doing all this, I have become aware of an entire set of risks to my privacy and security that, just a few weeks ago, were unknown unknowns for me. I thought it might be useful to share these observations, as well as some suggestions for addressing them. So I’m writing a series of posts that explains each one in plain language. It’s a work in progress.
Why do non-coders face so many unknown unknowns in Claude Code?
Perhaps surprisingly, the answer to this question came from a conversation I had with Claude Code itself. And it seemed obvious, once I learned it. As with every AI assistant, Claude Code is explicitly designed to reduce friction for the user. In other words, the whole point is to make things easier for you, the user, and to get you to a finished product faster. Some of the friction being removed is wonderful: I used Claude Code to build me a professional website in an afternoon, a task which would have taken me many months if I’d had to learn HTML and whatever other coding is going on behind the scenes. But some of the friction once did protective work: someone writing their own code is at least in a position to recognize when a command exposes information they’d prefer was kept private, and they can pause and make sure they decide to keep it that way. When that pause is engineered away to create a frictionless, goal-attainment machine, such exposures become invisible. The convenient path and the cautious path look identical to a non-coder Claude Code user, and a system optimized to be helpful has no particular reason to volunteer the difference.
The two common responses to Claude Code’s unknown unknowns
For someone in my position—curious, capable, pressed for time, and with essentially zero coding knowledge— the “unknown unknowns” situation can look like a choice between two doors, neither of them good.
Door Number One: abstain. You protect your privacy and security by declining to hand your work to a system you do not understand. This is the safe course, and not an unreasonable one. Its cost is everything the tools might have done for your research, your teaching, your institution—together with the possible experience of watching colleagues accelerate their work while you keep to your regular pace.
Door Number Two: use the tools blindly. You adopt them, enjoy the speed, and decline to think too hard about the machinery underneath. This is, in my observation, what many people have in fact done. I did it myself, at the start. But this is exactly where unknown unknowns can wreak havoc on privacy and security. Because non-coders have little experience that might lead us to raise the relevant questions, we are flying blind. We do not know what we do not know.
Door Number Three: Learn the unknown unknowns and address them
I’m never going to be a coder. Just to give you a sense of my non-coder-ness, consider that three different computer science postdocs tried to teach me the basics of GitHub (a version control system used by computer scientists and many others), and I failed to learn it each time. But thanks to the encouragement of Kemal Badur, the University of Chicago’s Chief Technology Officer, I did learn how to open the Terminal app on my Mac and talk with Claude Code on the Command Line Interface (CLI). I do wonder why the font in Terminal is terrible and so small. It’s been exciting to see what I can get out of those interactions with Claude Code, but my realizations about unknown unknowns had to be seeded by discussions with actual people. Through talking with BK Lee (a sociologist from New York University) and Nick Feamster (a computer scientist at the University of Chicago with whom I frequently collaborate), I’ve been able to understand certain things about how Claude Code operates. This led to a “eureka” moment about the existence of unknown unknowns for non-coders like me. With that insight, I asked Claude Code some questions, and it helped me come up with a taxonomy of what the big unknown unknowns about itself are. And so I asked it what I needed to do to protect myself from them.
From what I can tell, there are about six key unknown unknowns that non-coders should understand and take precautions against before they start using Claude Code in earnest. Some of what you’ll need to do to protect your privacy and security while working with Claude Code is a one-time setup task. Some of it will take you an afternoon to figure out. Others require ongoing discipline, but your awareness of them means they are no longer unknown unknowns. None of it requires you to become someone you are not.
Six unknown unknowns in Claude Code
I’m hoping I can produce the details on each of these six things in a fairly compressed window of time. Yes, Claude Code will help me do it. Here’s a preview of what they are:
What Claude Code, OpenAI’s Codex, or any other AI assistant can touch and do on your own computer, and how you can confine it
What you owe the other people whose data you hold—students, patients, clients, research subjects—before you enter any of it into an AI assistant
What data actually leaves your computer when an AI assistant reads or creates files, and how you can restrict it
Where that data goes, and how you can decide whether it is used to train future AI models
What is retained on your own computer once your AI assistant session ends
Whether you can, in fact, tell what an AI assistant is doing—this is the hardest one, and cuts across all the others
Each installment in the series will explain the issue (so the unknown unknown becomes known), then close with something you can do that same day to address it: a setting to change, a question to ask, an item for a checklist kept beside your keyboard.
My aim is not to usher you through Door Number One. Abstaining is safe, but its cost—everything these tools might do for your research and your teaching—is higher than it needs to be. Nor is it to abandon you at Door Number Two, where the benefits are real but the unknown unknowns are left free to do their damage. My goal is to show that the choice between those two is a false one. The risks that make blind use dangerous are neither endless nor mysterious. Door Number Three is simply Door Number Two with the unknowns made known. You use these tools in the ways that advance your work, while knowing where the dangers lie and how to guard against them. For those of us who are not computing people, and have no plans to become so, that is the door I am trying to hold open.