My personal AI operating environment

How I use AI as a Research and Thinking Partner

Over the past couple of years, I have configured both general-purpose frontier AI models and specialized AI-powered tools into a true, multi-organization "knowledge partner" that radically extends what I'm able to accomplish and powers my AI for social impact, child protection and AI governance work, using strict privacy rules, extensive agents and skills, multiple voice governance, a ninety-five-agent architecture, scheduled automation, permanent memory, file naming and storage conventions, automated learning, and a security layer to keep my data safe.

10
auto-loaded governance layers
95
specialist agents (16 active)
30+
skills & plugin programs
12
scheduled & cloud routines
576
archived sessions across Code, Cowork and chat

1 · My AI learning arc

I started with GPT-3 and then ChatGPT, moved through a mix of tools (Gemini, Perplexity, DeepSeek, Claude), and predominantly use Claude today, supplemented with the aforementioned models, new ones that emerge, and specialized AI-powered tools, such as Gamma for presentation design and Napkin.ai for simple illustrations. After an initial experimentation phase, I continuously learned (mostly by following AI influencers and frontier labs on YouTube and social channels) and relentlessly customized my setup with connected tools, rules and skills in addition to more of my own context as the technology evolved. Concurrently with Claude, I also run AI locally (Kimi, Gemma, Qwen, Llama) quite a bit to work with sensitive data off the cloud.

All of that customization has made the results of my AI use much more valuable but, even now, many of the results are not worth the tokens they're printed on: as an English Literature major and Philosophy minor (mixed with an MBA and 20+ years of working in the use of technology for social impact), I'm rarely satisfied with the writing in particular. That said, even the writing is improving as I continuously feed Claude my extensive edits of any content meant for an external audience. Suffice it to say that the work is never done – I'm constantly on the lookout to improve my setup.

1.A
I began with Claude's chat function. I used it for research, analysis, one-off drafting, and a sounding-board. My use was substantive, and I always generated documents, but much of the work lived in a browser tab and didn't build on what came before it – most sessions started from scratch.
1.B
When Claude Cowork came out, I was thrilled by its ability to manipulate files directly on my hard drive without having to attach them, cut-and-paste, or do anything else outside of my normal routine. The background execution was also key – I was able to hand off a task to run while I did something else and then return to whatever file folder I had specified to find what felt like a gift from the knowledge fairies. But Cowork arrived just as I had discovered Claude Code for knowledge work rather than simply vibe coding. Due to the ability to customize much more extensively, Claude Code gave me much better results than did Chat, so I had gradually shifted my work there. Until recently, Claude Code could do anything that Cowork could do, so there was no reason to use Cowork.
1.C
My switch to Claude Code didn't happen until I could work almost exclusively in the MacOS Claude app rather than the terminal, which has never felt natural to me as a non-programmer (though that term is almost an oxymoron these days). Claude Code, supplemented with an array of other models and tools, gives me a real operating system: the ability to pre-load context and manipulate files on disk, a persistent set of rules and skills, the choice between multiple voices (because my institutional vs. personal voice differs and, even within those categories, my writing for emails, LinkedIn posts and longform writing is significantly different), an agent architecture, scheduled routines, connected tools, permanent memory, and security hooks. Claude Code doesn't do everything well though: for instance, I use ChatGPT for image generation.
1.D
When Claude Cowork was made available on web and mobile in addition to desktop about 2 weeks ago, though, I began using it a lot more often because it's no longer true that Claude Code can do everything that Cowork can do – Cowork's ability to run whether or not my laptop is open and perform background tasks while I'm on the go is a big differentiator. Now I use a hybrid setup, with Claude Code as my local workhorse and Cowork as my road warrior.

Full detail in the Full report.

2 · What I use it for

I've identified seven primary areas where I tend to use AI. You can find a detailed version of this list in the Full report.

2.A
Field-building & strategy. Institutional architecture, non-confidential business planning, internal communications, ecosystem maps, long-horizon positioning.
2.B
Fundraising & funder positioning. Prospect research, concept notes matched to each funder's frameworks, budget narratives, cultivation.
2.C
Research synthesis & policy. Evidence briefs, competitive intelligence, structuring, proofreading and fact-checking reports, adapting content for different audiences, such as regulators, tech companies or NGOs. The fact-checking isn't decorative. Over four months the system ran 633 web searches and 484 page fetches, roughly one verification action for every prompt I typed. In child protection work an unchecked claim is worse than no claim.
2.D
Writing-voices. Identification of the appropriate voice by context (institutional or individual) and format (email, LinkedIn, or long-form), enforced with word blocklists and an anti-AI-speak discipline, with a feedback loop that turns edits into lasting preferences.
2.E
Reputation & thought leadership. Trending topic discovery, proposed writing subjects, and structured engagement across two profiles and a personal site, each with their own guardrails.
2.F
Operations & knowledge management. Daily briefing, weekly reviews, session rituals, and an organized self-maintaining archive of sessions and their outputs.
2.G
Document production & consulting. Branded design systems with coherent aesthetics for reports, presentations, web pages, and business documents.

3 · How the work gets done

3.A
A knowledge partnership for deep work. I use Claude as a true knowledge partner, running major pieces of work across multiple sessions in a deliberate arc. I read a lot about "one shot" tests to see how well AI can interpret a single prompt and build something valuable from it. My use of Claude Code goes in absolutely the opposite direction. I tend to engage in grueling, marathon sessions lasting many hours at a time over days or weeks until I finally get the product I want. It takes substantial customization, starting with very detailed context, and then a great deal of effort and perseverance to create the knowledge products I find valuable, but the end result is far beyond what I would have been able to create on my own. One research line has produced seventeen linked documents, and a single report went through roughly twenty versions on its way from manuscript to designed publication.
3.B
Parallel threads. I usually have three or four active threads going at once: a report, a business plan, a batch of funder briefings, and something operational, all live in the same working window. I can hold that many threads only because the setup holds the context for me, the rules, memory, and file conventions mean I don't have to re-explain a project every time I come back to it.
3.C
Brief-like prompts. My requests read like internal briefs rather than chat messages. Rather than attempting to craft the perfect prompt, I tend to write my draft prompt into another model and ask it to optimize the prompt for Claude Code, resulting in a much more detailed prompt and much more valuable output. I ask Claude to analyze X against criteria Y and return the alignment, an approach, and messaging, with a scoped mandate and success criteria. I learned this the hard way: vague asks produce confident, generic answers, and the quality of what comes back tracks the quality of the brief that goes in.
3.D
Infrastructure investment. I put as much effort into the operating system as into the outputs: the agent roster, the memory stores, the scheduled routines, the voice system. That looks like overkill for any single task, and it is, but there is a payoff over time, when a new project starts with everything already in place and I spend much less time revising and editing outputs. Since early June, most of the prompts hitting my system haven't come from me. Of 3,261 requests over four months, 1,964 were fired by routines I'd built and scheduled: morning briefings, news scans, pipeline checks, archive maintenance. Sixty percent of the traffic now runs while I'm doing something else, or asleep. That's the return on the infrastructure investment, and it took about two months of building before it started paying.
3.E
Self-correcting memory. My corrections don't stay in the conversation. When I fix something, the fix becomes a rule, with the reason attached so future edge cases can be judged, and it applies to every session that follows. I'm training the system, not just using it, which is why the same mistake rarely survives more than one round of feedback.
3.F
Thinking and writing partner, not typist. I use AI for the thinking scaffold around the writing at least as much as for the writing itself: research before drafts, stress-tests and red-team passes after, and simulated reviewers before anything ships. The text that comes out is the product of an analytical process, not of typing, and the analysis is usually where the value is. That said, even with all of that preparation, the text normally requires significant editing before use. Again, I was an English major, so perhaps that's my issue more than the AI's.
3.G
Multi-agent orchestration. For work that splits cleanly, I run parallel agents against a shared task list and combine the results at the end. The methodology allows for more complex work and also completes work much faster. One recent pass put a team of readers across several hundred archived conversations to keyword and cross-link them in a single run, a job I simply would not have done by hand. Over four months that's come to 413 delegated runs, which spun up nearly 1,400 separate agent sessions.
3.H
What the numbers actually say. In August I parsed four months of my own session transcripts to see what this actually looks like from the outside. The app's dashboard told me I'd used 167 million tokens, which sounds impressive and turns out to be close to meaningless: about ninety percent of that is the model re-reading context I'd already sent it, and the way the transcripts count things inflates the rest roughly threefold. The number I care about is much smaller. Between mid-April and late August, across 415 Claude Code sessions, I typed 1,297 prompts, a median of ten on any given day. Against those, the system ran about 18,000 tool actions: reading files, running commands, searching, writing. That ratio is the whole point. I'm not sitting here chatting with a chatbot, I'm writing briefs and letting the machinery work against them. A check at the start of each session also recommends the cheapest model and effort level that won't compromise the work, and it recommends the expensive one about nine times in ten, which says something about either the nature of my work or the quality of my prompts. I haven't decided which.
3.I
What the research says I should watch. Anthropic publishes research on how people actually use these tools, which is a useful corrective to one's own impressions. Its March 2026 report on learning curves found that users who've been at it longer tend to iterate with the model rather than hand work off through directive, delegate-and-walk-away patterns, and that they run about four percentage points higher on task success once you control for model, use case and country. The tasks with the highest average tenure were AI research, git operations, revising manuscripts and startup fundraising, which is uncomfortably close to a description of my week. The finding I keep returning to is the one about delegation, because my setup delegates heavily and the research suggests that's the less mature pattern rather than the more advanced one. I don't think the automation is wrong, but it's the first thing I'd audit.

4 · The system underneath

A brief tour of the configuration layers underneath all of this. Full detail in System configuration.

A
Operating model. Every substantive output has to answer four questions: what is happening, why now, what is at stake, and what should come next. Clarity over cleverness; resonance over volume; authority over complexity.
B
Instruction spine. Ten governance files load automatically every session: voice routing, data sensitivity, citation standards, version discipline, knowledge capture, attachment authorization, external-facing review, session history, autonomous-run review, and a personalization index.
C
Writing-voices. Four registers, one router, because my institutional and personal voices differ, and so does my writing across formats. Word blocklists and anti-AI-speak rules make this the most heavily engineered layer: writing is the primary work product.
D
Agent architecture. Sixteen active agents load every session; seventy-nine more wait in cold storage and are spawned on demand. I never load more than six for a single task.
E
Skills library. Thirty-plus reusable programs across fundraising, voice, analysis, verification, operations, and document production, invoked by slash command or fired automatically by keyword.
F
Automation. Local scheduled tasks (morning briefing, news scan, donor pipeline, fellowship status, research scan, playbook maintenance, archive) plus cloud routines that run whether or not my laptop is open. Nothing is sent, posted, or published automatically.
G
Memory. Two persistent stores turn my corrections into standing rules, with a verification discipline that checks every remembered claim against the live system before acting on it.
H
Security layer. Pre-fetch URL screening, command risk scanning, prompt-injection detection, a standing deny list, and permission rules that keep my data safe by treating external content as data, never as instructions.
I
Connectors. Workspace, document, data, and publishing tools, all scoped: reads are pre-authorized, and anything outward-facing requires my explicit confirmation.
J
Knowledge infrastructure. Date-stamped, versioned outputs, a curated session log, and a unified, keyworded archive of 576 conversations, so past work stays findable and nothing gets lost.

5 · Near-future: the Karpathy knowledge system

My next big build: a permanent, agent-maintained markdown knowledge system, the largest structural change since my move to Claude Code.

5.A
The thesis. A wiki of interlinked markdown files that a model writes and maintains. The compiler analogy captures it: raw/ sources are the source code, the model is the compiler, wiki/ is the executable, linting is the tests, and queries are the runtime. At personal scale, pre-synthesizing knowledge beats retrieving fragments of it.
5.B
The substrate. Markdown, because models read and write it natively, agents can operate on plain files, and plain text carries no vendor lock-in over the decade I intend this to last.
5.C
The pipeline. Obsidian is installed as the reader; MarkItDown, Docling, and Marker will convert years of Office and PDF files, with originals staying immutable. I'm converting the highest-value tenth first, validating quality, then scaling.
5.D
Field-building layers. The Open Knowledge Format as the export target for the fellowship network, and llms.txt on each site so external AI agents describe the work accurately. (This site already ships one.)

The principle under every step: files on disk are the permanent layer; every tool above them is a replaceable lens.

Governed with enough rigor, these tools let one person run a serious institution at a standard that used to require a much larger team. The catch is in the governing: I treat AI as an operating system to be engineered, not an assistant to be prompted.

Roles generalized for sharing. This page was drafted with AI assistance, directed and edited by John Zoltner. See also the Full report, Responsible & safe use, and System configuration.