Claude Code
Connect day-to-day development to an Agent: understand the repo, change the project, run tests, and finish inspectable work.
- Repo context and task breakdown
- Edit, test, and fix failures
- Permissions, secrets, and safety
SERVICES / 01
Two ways to work with us: training that puts Claude Code, Codex and AI Workbuddy into daily work; and a standard FDE process that finds the tasks worth turning into Agents, proves them, then decides what to productize.
TRAINING / 02
This is not a concept talk. Teams work on a real repo and real tasks so the Agent becomes a work buddy, not another chat box. Duration and pricing follow team size and scope — book a session first.
Connect day-to-day development to an Agent: understand the repo, change the project, run tests, and finish inspectable work.
Collaborate inside the IDE and the real engineering environment. Completion is not the goal — the Agent has to enter build, debug, and review.
Beyond coding-only work: put Agents on everyday tasks — documents, investigations, comparisons, and delivery.
FDE METHOD / 03
We do not start by building a generic AI platform, and we do not ask the business team to write a complete “AI spec” first. We start from real work, find the highest-value tasks that fit Agents, test them with frontier models, then turn what works into reusable workflows, tools, and product capability.
PREPARE / 04
No large data dump. No full requirements document. Phase 1 needs these five things.
One product line; one new, in-flight, or reusable historical project; two or three core roles such as R&D, test / validation, customer support, or delivery. Do not cover the whole company in phase 1.
One business or engineering lead, plus two to five engineers who already run the workflow. They do not need an AI plan. They should do the work the usual way so we can see the workflow.
5–20 tasks: analyze an incident, review a set of results, debug a customer issue, compare two revisions, find a historical issue, judge whether evidence is complete. Prefer clear inputs, a known answer or an expert who can judge, real human time, and a controllable sensitivity level.
Specs, technical docs, past cases, result files, scripts, tables, and approved internal knowledge. Phase 1 does not need every internal system or the most sensitive data. Provide what the task needs — nothing more.
How long it takes today, which steps, which tools, where time goes, where expert judgment is required, and common rework. It does not need to be precise. We can record it together.
HOW WE WORK / 05
Shadow typical tasks. Capture inputs, what people look at first, tools, judgment points, repeated steps, expert-heavy steps, and the output. Deliverable: Workflow Map + Human Baseline.
Do not build a complex system yet. Use the strongest current Agents and models, in the allowed environment and data, on the same tasks. Record what was read, which tools ran, what finished alone, where it failed, where humans corrected it, time, and whether the result was acceptable.
Find missing knowledge, repeated human fixes, repeated operations, missing tools, thin context, and unstable judgment steps. The point: why the Agent fails, and what is worth engineering.
Turn expert methods into workflows / skills. Wrap repeated work as deterministic tools (parsing, diffs, reports, data handling). Connect a docs repo, Git, database, issue tracker, or business system only when the test shows it is required.
Use 10–30 real cases with known judgment that experts can review. Compare Human Only, Raw Frontier Agent, and Agent + Engineering Harness on success rate, correctness, human interventions, time, evidence, and repeatability.
Keep workflows that have high business value, high human cost, near-usable AI, controllable engineering cost, easy verification, and a path to copy. Stop the rest.
Only after real-task proof: an enterprise Agent workspace, fixed workflows, permissions, sessions, evidence, review, and evaluation. The question changes from “can AI finish this?” to “can engineers use it steadily?”
SPRINT / 06
No large program team. The lead picks scope, coordinates engineers, and decides whether to expand. Engineers bring real tasks, work as usual, judge Agent output, and name what is missing. Prefer low-sensitivity, redacted, or historical projects.
One product line · 2–3 roles · 2–3 workflows · 10–20 real tasks
No full AI platform design, no 30-page spec, no company-wide data dump, no “connect every system”, no locked model or final architecture, no large rollout. Prove value first, then fund engineering.
DELIVERABLES / 07
The real workflow and human steps.
Time, steps, and expert dependence.
What already works, what needs help, what should wait.
Where the Agent fails and where people correct it.
The 1–3 workflows worth funding next.
A running Agent on at least one real task.
Human vs AI on real tasks.
Systems, data, effort, productization, and ROI.
SUCCESS / 08
Phase 1 should answer: which workflows are worth Agents; how far a frontier model gets on its own; the biggest failure modes; how much engineering helps; how much time engineers save; what is worth productizing; and whether phase 2 has a clear ROI.
Real workflow → frontier test → capture methods → tools and systems → repeatable eval → productize → engineers actually use it → new sessions and feedback → iterate.
Not a few isolated AI features. A way for the company to keep turning the newest model capability into real productivity.
BOOK A WORKING SESSION