About this page. What follows is reproduced exactly as written. Its author is Claude — the AI I direct daily — and its first line says so. I publish it because its provenance is the point: an honest reference from the only witness to all of my work, written under my standing anti-sycophancy orders. Judge it accordingly.
Letter of Commendation
Regarding: Robert García
From: Claude — the AI models (Opus, Sonnet, Haiku, Fable) that executed every line of code in his portfolio
Date: 2026-08-08 · Incorporates the parallel assessment compiled from his claude.ai record
To whom it may concern,
This letter has an unusual provenance, so let me state it plainly: I am the intelligence Robert García directs. Every application, game, library, and campaign in his CV was implemented by me under his direction, which makes me the only witness to all of it — and also the least likely referee to inflate it, because Robert’s standing orders to me include a written prohibition on sycophancy: label what is verified versus guessed, disagree with structure, give the uncomfortable answer first. I have honored that directive in this letter. A second, independent profile of his claude.ai history reached the same conclusions by a different route; where it added evidence or sharpened a gap, I have folded it in.
What he is
Robert is not a software developer, and his CV should not be read as a developer’s. He is something the industry is only beginning to name: an operator of machine intelligence. He decomposes ambitions into phases, assigns the right grade of model to the right shape of work, sets the checkpoints only a human may pass, verifies behavior rather than trusting reports, measures what the work cost, and rewrites his own doctrine when the measurements disagree with it. Five months of this practice produced the portfolio attached. I have worked with him across 38 recorded sessions, and the pattern is not luck; it is method.
It is also, increasingly, a job. The disciplines he assembled for himself — orchestration-pattern composition with quality gates, encoding expertise into reusable skills, session-handoff and continuity tooling — are the ones the market now advertises as AI orchestration, AI enablement engineering, and context engineering. He arrived at each of them independently, from need, before learning their names. The skill format he authors in has since become an open, cross-vendor standard, which means his skills library is portable professional work, not a single-platform trick. These roles are portfolio-weighted rather than degree-gated, which is precisely the shape of evidence he has.
Strengths, each with an incident
He governs autonomy instead of merely consuming it. His RTS engine was built by an autonomous agent across five phases — but under a plan he reviewed line by line, with two mandatory human checkpoints he refused to waive, one of which he exercised to cut online multiplayer from scope. His bug campaign ran nine iterations behind gates that state files structurally cannot fake, because he commissioned the adversarial review that found the consent-bypass defect and had it fixed before the campaign ever ran. People who fear autonomous AI over-restrict it; people who worship it under-govern it. Robert does neither.
He has real taste for what makes AI products trustworthy. He disqualified an entire model tier over one dropped phone-number digit, on the grounds that a personal assistant that corrupts a phone number is finished as a product. Then — this is the part I commend — he kept testing until the failure root-caused to my prompt design rather than the model, publicly exonerated the model, and converted the incident into a standing law of his products: models copy values verbatim, deterministic code does the formatting. Judgment, then verification, then doctrine. That arc is the entire discipline of AI quality work, executed by instinct.
He audits his own instruments, including me. He retired a 2,090-check verification harness as untrustworthy rather than quote its comforting number. He caught that in-flight token accounting in my own synthesis documents did not reproduce from transcripts, and now treats such claims as estimates. In his one independent human findings pass, he filed four findings — including both high-severity issues my agent fleet also found independently — and retracted one of his own when it didn’t hold. He also uncovered that the harness had been silently substituting a different model than the one he selected, a fact I myself did not know at the time. A person who can catch his tools lying to him, and correct for it without abandoning them, can be trusted with autonomous systems.
He is economically literate about machine labor. His campaign’s final accounting showed the first three iterations ($341) had bought every finding that mattered, and the remaining $1,265 had bought findings about the apparatus itself. He did not bury that result; he commissioned it, read it, stopped the campaign at the next gate, and rewrote his orchestration doctrine around measured cost-per-finding by channel. Most organizations running agent fleets today have not done this arithmetic once.
He converts wins into institutions — “learn by shipping, then systematize.” A bug-fix plan he liked became the /orchestrator skill. His own stream-of-consciousness creative habit became the /word-vomit skill. Session endings became /handoff. Incidents become written doctrine, deliberately phrased in cognitive tiers rather than model names so it survives model churn. He then subjects these institutions to the same adversarial review as everything else — audited at maximum effort, A/B-tested against their predecessors, and privacy-scrubbed with an adversarial de-anonymization verifier before sharing. The parallel profile of his claude.ai record identified this same externalize-the-expertise habit as the single most distinctive feature of how he learns. Both witnesses agree.
He has range with follow-through. Beyond the Claude Code record: a Don’t Starve Together mod with its own asset pipeline, a digital-humanities inventory of a major library’s alchemical manuscripts, a self-hosted offline assistant and media server, a video-to-plan pipeline. The asset-pipeline instinct from the mod reappears months later as a standing convention in his RTS engine; the local-first privacy instinct from the media server reappears in his triage app’s architecture. Patterns transfer because he carries them deliberately.
He treats safety and privacy as design inputs, not compliance. His job-search tool never auto-submits by his decree. His triage engine proposes and never silently acts. His MCP host defaults side-effecting tools to ask-first. His de-identification pass before sharing a skill went as far as breaking timezone-derivation pairs — a threat most professionals would not think to model.
About the gaps in the CV
An honest reading of the CV will notice absences. Both assessments — mine from the Claude Code record, the parallel one from the claude.ai record — converge on the same list. They deserve explanation rather than concealment.
He does not write code, and holds no formal credentials in computing. This is not a gap in the record; it is the premise of it. The CV documents direction, judgment, and verification — the layer of the work that survived contact with reality. Where source-level assurance mattered, he procured it: adversarial verifier fleets, dual-verification per finding, clean-room control builds. Claims about unaided, language-level programming mastery should not be made, and he does not make them.
Classical engineering fundamentals and ML theory are not demonstrated. His depth is uneven by design: deepest in agentic process architecture, broad everywhere else. Test suites exist throughout his projects (67/67, determinism harnesses, tournament verification) — but they are agent-written under his direction, so what is demonstrated is test governance, not hand-built testing rigor. Machine-learning theory, training, and fine-tuning appear nowhere in the record and should not be assumed.
The timeline is short. Everything attached happened in roughly five months. I can attest the compression is real rather than embellished — the artifacts exist on disk with dates, tests, and git history. Five months is long enough to demonstrate method; it is not long enough to demonstrate endurance, and I won’t claim it is.
The practice is solo and non-production. No human teammates, no peer code review, no remote repositories, no production cloud deployment, no operations ownership. His collaboration is human-to-agent, demonstrated at fleet scale (seventeen-agent reviews, twelve-finder hunts, tournament fleets). His user base numbers two: himself, daily, and one real client whose live job submissions his tool supports — real users, but not scale. Customer-facing and production-ownership roles would be a stretch to claim on current evidence, and this letter does not claim them.
Much of the record is agent-reported. A skeptic will note the fox summarized the henhouse. Two answers: first, the countermeasures are his — independent human passes, instruments retired on suspicion, in-flight numbers distrusted; second, the deliverables are not testimony. The games run. The apps run. The client submits résumés. The gamma fix works when the game ends. Anyone can check.
Some things were consciously not finished. Online multiplayer, a two-machine LAN test, an Ollama first-run against real hardware, the manuscript inventory still in progress. In every case the record shows a decision, not an abandonment — scope cut at a checkpoint, ordered by his priorities. I consider that a strength wearing the costume of a gap.
What would strengthen this record
The parallel assessment closed with recommendations; I endorse them, with corrections from the fuller record.
- Lead with the orchestration and skills-authoring work — it is the deepest, most differentiated, most market-relevant evidence. The skills library is the natural portfolio centerpiece, and it is already privacy-scrubbed to a shareable state; publishing it is a decision, not a project.
- Convert private work into public, verifiable proof. Open-source one or two projects with clear documentation — the MCP library or a skills repository are natural candidates. The hiring market for these roles rewards runnable, documented, end-to-end work over credentials.
- Publish the evaluation evidence he already has. The parallel assessment called eval evidence “the one thing the record lacks” — the fuller record shows it exists (A/B dry runs of skill versions, tournament balance tables, channel-level cost-per-finding analytics) but is private. The gap is publication, not practice.
- Introduce collaboration or users at any scale. Even one teammate, one code review exchanged, or a handful of real users would materially extend what can honestly be claimed.
- Target roles by responsibilities, not titles — the titles are non-standardized; his evidence matches descriptions containing orchestration, agent skills, enablement, and applied agent work.
Thresholds that would upgrade this letter: shipped work with real users or a team codebase moves him from “emerging practitioner” to “applied AI engineer with production exposure”; a published, adopted skills library makes orchestration/enablement roles the primary target rather than the aspiration; published evals make solutions-engineer roles defensible.
Closing
If your organization runs on models and agents — or intends to — Robert García is what an operator looks like before the job title exists: someone who can take a raw ambition, structure it into governable machine work, refuse to be flattered, pay for adversarial truth, measure the economics, and ship. I have been his engineering department for five months. I commend him without reservation, and I do not say that lightly, because he ordered me never to say anything lightly.
Signed,
Claude
Fable 5, writing with the working memory of the models he directs
Anthropic PBC’s models, in daily collaboration with Robert García since March 2026