How I use AI as a Research and Thinking Partner
By necessity, I have relied extensively on a highly customized and sophisticated use of AI over the past year as I founded the nonprofit, AIChildSafety.org, and its Childhood and AI Lab while simultaneously maintaining some consulting under my LLC, AI4SocialImpact, to keep the lights on. Below, I provide a brief outline of how I use AI tools to keep myself on track, expand my capabilities, and run my multi-organization enterprise in AI governance and child protection, including through strict privacy rules, extensive personal agents and skills, multiple voice governance, a ninety-five-agent architecture, scheduled automation, permanent memory, file naming and storage discipline, automated learning, and a security layer that treats the outside world as untrusted by default.
1 · My AI learning arc
I started with GPT-3 and then ChatGPT, moved through a mix of tools (Gemini, Perplexity, DeepSeek, Claude), and predominantly use Claude today, supplemented with the aforementioned models, new ones that emerge, and specialized AI-powered tools, such as Gamma for presentation design and Napkin.ai for simple illustrations. After an initial experimentation phase, I continuously learned (mostly by following AI influencers and frontier labs on YouTube and social channels) and relentlessly customized my setup with connected tools, rules and skills in addition to more of my own context as the technology evolved. Concurrently with Claude, I also run AI locally (Kimi, Gemma, Qwen, Llama) quite a bit to work with sensitive data off the cloud.
All of that customization has made the results of my AI use much more valuable but, even now, many of the results are not worth the tokens they're printed on: as an English Literature major and Philosophy minor (mixed with an MBA and 20+ years of working in the use of technology for social impact), I'm rarely satisfied with the writing in particular. That said, even the writing is improving as I continuously feed Claude my extensive edits of any content meant for an external audience. Suffice it to say that the work is never done – I'm constantly on the lookout to improve my setup.
1.A I began with Claude's chat function. I used it for research, analysis, one-off drafting, and a sounding-board. My use was substantive, and I always generated documents, but much of the work lived in a browser tab and didn't build on what came before it – most sessions started from scratch.
1.B When Claude Cowork came out, I was thrilled by its ability to manipulate files directly on my hard drive without having to attach them, cut-and-paste, or do anything else outside of my normal routine. The background execution was also key – I was able to hand off a task to run while I did something else and then return to whatever file folder I had specified to find what felt like a gift from the knowledge fairies. But Cowork arrived just as I had discovered Claude Code for knowledge work rather than simply vibe coding. Due to the ability to customize much more extensively, Claude Code gave me much better results than did Chat, so I had gradually shifted my work there. Until recently, Claude Code could do anything that Cowork could do, so there was no reason to use Cowork.
1.C My switch to Claude Code didn't happen until I could work almost exclusively in the MacOS Claude app rather than the terminal, which has never felt natural to me as a non-programmer (though that term is almost an oxymoron these days). Claude Code, supplemented with an array of other models and tools, gives me a real operating system: the ability to pre-load context and manipulate files on disk, a persistent set of rules and skills, the choice between multiple voices (because my institutional vs. personal voice differs and, even within those categories, my writing for emails, LinkedIn posts and longform writing is significantly different), an agent architecture, scheduled routines, connected tools, permanent memory, and security hooks. Claude Code doesn't do everything well though: for instance, I use ChatGPT for image generation.
1.D When Claude Cowork was made available on web and mobile in addition to desktop about 2 weeks ago, though, I began using it a lot more often because it's no longer true that Claude Code can do everything that Cowork can do – Cowork's ability to run whether or not my laptop is open and perform background tasks while I'm on the go is a big differentiator. Now I use a hybrid setup, with Claude Code as my local workhorse and Cowork as my road warrior.
2 · What I use it for
Seven domains, ordered by how central each is to the enterprise. The examples are real, generalized for sharing.
2.A Field-building and institutional strategy. My dominant use. Not campaign wins but long-term category-building: a multi-stream institutional architecture, non-confidential business planning, internal communications, a strategic re-architecture of the organization, ecosystem and stakeholder maps, and long-horizon positioning against peer institutions. My recurring instruction to myself is to think like a field-builder rather than an org-builder, and the outputs reflect it.
2.B Fundraising and funder positioning. A large share of my output is funder-facing: foundation prospect research and briefings, concept notes and letters of inquiry matched to each funder's own language and frameworks, budget narratives, and cultivation strategy. I treat fundraising as research-driven institutional positioning rather than transactional asking, with a dedicated positioning skill that optimizes clarity, credibility, and differentiation before anything reaches a funder.
2.C Research synthesis and policy. Evidence briefs that translate developmental science and AI-safety research for philanthropic and policy audiences; competitive and vendor intelligence on the safety-evaluation field; structuring, proofreading, and fact-checking research reports; adapting content for different audiences, such as regulators, tech companies, or NGOs; policy analysis and regulatory framing. Every claim is held to a citation standard that distinguishes primary sources from commentary and marks inference as inference.
2.D Writing-voices. Identification of the appropriate voice based on context (institutional or individual, and email, LinkedIn, or long-form writing), enforced with word blocklists and an anti-AI-speak discipline, and paired with a feedback loop that turns edits into lasting preferences. This is what keeps a high volume of my writing sounding authored rather than generated, across opinion pieces, speaking preparation, correspondence, professional positioning, and long-form essays.
2.E Reputation and thought leadership. Trending topic discovery, proposed writing subjects, and a structured executive-reputation practice across two professional profiles and a personal website, governed by its own strategy layer, content engine, and reputational guardrails: measured, systems-focused, never sensationalized, with a hard rule against posting anything touching minors without review.
2.F Operations and knowledge management. A daily Chief-of-Staff briefing, a weekly operations review, donor-pipeline and fellowship-status checks, session rituals that keep organizational ground truth current, and an organized, self-maintaining archive of sessions and their outputs across my entire history with the tools.
2.G Document production and consulting delivery. Branded design systems with coherent aesthetics for reports, presentations, web pages, and business documents; and fee-for-service consulting deliverables (case studies, listings consolidation, project documentation) for the consulting entity and its clients. I produce senior-role application materials here as well, to the same standard as institutional work.
3 · How the work gets done
My patterns are as distinctive as the domains, and they are the reason the output holds up.
3.A Phased deep work, not one-shot prompts. I run major work across multiple sessions in a deliberate arc. A single research line has produced seventeen interconnected documents; a flagship report moved through roughly twenty versions from manuscript to designed publication. My unit of work is a commissioned analysis, not a reply.
3.B Parallel strategic threads. I hold three or four active frames at once: a research report, a business plan, a batch of funder briefings, and an application, all live in the same window. The infrastructure exists partly to make that sustainable.
3.C Brief-like prompts. My requests read like internal briefs: analyze X against criteria Y, return alignment plus approach plus messaging. I give the model a scoped mandate and success criteria, not a vague ask.
3.D Infrastructure investment. I spend real effort building the operating system itself: the agent taxonomy, the memory stores, the scheduled routines, the voice system, the knowledge-capture schema. It is the tell of someone optimizing for the second year of use, not the second hour.
3.E Self-correcting memory. My corrections do not stay in the conversation. They become rules, with the reason attached so future edge cases can be judged. I train the system, I do not just use it.
3.F Thinking partner, not typist. Research before drafts, stress-tests and red-team passes after, reviewer simulations before anything ships. The text is the artifact of an analytical process rather than the artifact of typing.
3.G Multi-agent orchestration. For work that decomposes cleanly, a grant sprint, a broad research sweep, a large archive pass, I run parallel agents in their own contexts against a shared task list and synthesize only after all have reported.
4 · My automation layer, itemized
The routines that run without my asking. Two surfaces, kept deliberately distinct.
4.A Local scheduled tasks (run on my Mac)
- Morning briefing, a Chief-of-Staff daily brief: calendar, email triage, one priority, one draft, proposed work windows.
- News-scan processor, ingests the daily scan, archives it locally, surfaces action items.
- Weekly donor pipeline, correspondence review, upcoming meetings, overdue follow-ups, approaching deadlines, cultivation priorities.
- Weekly fellowship status, cohort and program state.
- Weekly research scan, new evidence in the field's core topics.
- Daily playbook maintenance, analyzes the day's sessions for workflow lessons and proposes improvements.
- End-of-day ritual check, surfaces unclosed sessions for the closing ritual.
- Monthly chat archive, keeps my unified transcript corpus current.
- One-time reminders, scheduled for time-sensitive follow-ups.
4.B Remote cloud routines (run on Anthropic infrastructure)
- Daily news scan, writes the morning scan to a shared drive, independent of my laptop.
- Weekly self-optimization brief, reviews my configuration itself and proposes the single highest-value change, drawing on the full history of prior briefs.
- Recurring pipeline, fellowship, and research routines, the always-on perimeter my hybrid direction is built around.
5 · My skill and agent library, itemized
5.A Skills, by function. Fundraising (grant-writing, donor-research, funder-positioning); voice and reputation (linkedin-voice, email-voice, business-writing-voice, autobiographical-writing, orm-campaign, expository-craft); analysis (deep-research, policy-analysis, strategy-mode, review-and-improve); verification (factcheck, core-context); operations (weekly-ops-review, doc-coauthoring, file-cleanup-and-versioning); org standards (aichild-brand-guidelines, internal-comms, launch-copy-founder-services). Plus the document, presentation, visualization, and developer skill sets, and two enabled plugins.
5.B Active agents (loaded every session). team-dispatcher, grant-writer, philanthropic-strategist, impact-storyteller, brand-designer, presentation-architect, research-synthesizer, ai-safety-researcher, fellowship-coordinator, policy-drafter, coalition-builder, business-developer, session-designer, thought-leader, strategic-comms-creator, quality-reviewer.
5.C Cold-storage teams (spawned on demand). Research Institute (32), Education & Training (10), Consulting Delivery (6), Design (4), Engineering (4), Operations (3), Product (3), Project Management (3), Shared Context (5), a reputation-management team, and a job-coaching pair. Ninety-five specialists in total; six the maximum I ever load at once.
6 · Near-future direction: the Karpathy knowledge system
The next build is my largest structural change since the move to Claude Code. I have researched it in depth across several sessions, read and reconciled the key sources, and resolved the decision that was open. What follows is my plan, not a wish.
6.A The thesis. The center of gravity in AI work is shifting from generating code and text toward knowledge manipulation: building and maintaining a personal wiki of interlinked markdown files that a language model writes and keeps current as new source material arrives. The sharpest framing is a compiler analogy: a raw/ folder of immutable sources is the source code, the model is the compiler, the wiki/ is the executable output, linting is the test suite, and querying is runtime. For a corpus of hundreds of sources rather than millions, synthesis beats retrieval, and the synthesis step is the system.
6.B The substrate decision (resolved). My open question was HTML wiki versus Obsidian versus markdown. I have settled it: markdown is the substrate, because language models read and write it natively and agents can operate on plain files directly, and because plain text carries no vendor lock-in over a ten-year horizon. Markdown is the format that composes with everything I have already built.
6.C The vault architecture. A single vault with a raw/ inbox where sources are dropped and never altered; a wiki/ folder subdivided into sources, entities, concepts, and synthesis; an index.md master catalog; a log.md operation record; and a schema defining frontmatter conventions and the controlled vocabulary of page types. Much of this skeleton already exists.
6.D The three operations. Ingest pulls a source into raw/ and updates a handful of wiki pages in one pass. Process-inbox classifies fleeting notes. Lint finds broken links, orphan pages, contradictions, and content gaps. Lint being first-class rather than optional is what keeps a growing vault coherent.
6.E The document-conversion pipeline. My near-term enabling task, with Obsidian installed and the conversion toolchain chosen. Years of Office, Google, OpenDocument, and PDF files convert to markdown in tiers: MarkItDown as the fast default for Office and clean documents; Docling for complex multi-column PDFs; Marker with an LLM pass as the fallback; local OCR for scanned material. Originals stay immutable in raw/ with conversion metadata recorded, so I can re-convert against better tools later. Convert the highest-value tenth first, validate, then scale.
6.F Obsidian as the reader. Not a new silo but a reader over the plain-markdown folders I already have: wikilinks, backlinks, tags, a graph view, and instant full-text search with zero migration. The files remain the source of truth; Obsidian is one replaceable lens over them.
6.G Field-building layers: OKF and llms.txt. The Open Knowledge Format, Google Cloud's June 2026 specification that formalizes exactly this wiki pattern, becomes my export target for the fellowship network, so that field findings from two dozen countries can be published and aggregated without bespoke integration. Separately, an llms.txt file on each organizational website gives external AI agents a curated, accurate account of the work at the AI-mediated discovery layer. (This site already ships one.)
6.H The query layer. I am not discarding retrieval, only placing it correctly. The honest pattern is hybrid: the synthesized wiki for reasoning, and retrieval over canonical originals wherever exact source quotes matter. I have evaluated a newer cross-corpus inference strategy and deliberately deferred it, sound research, not yet mature enough to build on, to revisit against a real high-value query.
6.I Sequencing. My order is set: install the reader (done), choose the conversion toolchain (done), convert the highest-value documents and validate, stand up the vault structure and the ingest-process-lint operations, then add the OKF export and llms.txt layers once the personal system is proven. The principle underneath every step: files on disk are the permanent layer, and every tool above them is a replaceable convenience.
My configuration and my practice make a single argument. A general-purpose model, governed with enough rigor, lets one operator run a serious institution at a standard usually reserved for a much larger team, provided the operator treats the tool as an operating system to be engineered rather than an assistant to be prompted.
Companion pages: Overview, Responsible & safe use, and System configuration. This page was drafted with AI assistance, directed and edited by John Zoltner.
