SERVICES / 01

Start from a real workflow, not a generic AI platform.

Two ways to work with us: training that puts Claude Code, Codex and AI Workbuddy into daily work; and a standard FDE process that finds the tasks worth turning into Agents, proves them, then decides what to productize.

Real workflow Capability test Engineering Evaluation Productization

TRAINING / 02

Claude Code, Codex and AI Workbuddy

This is not a concept talk. Teams work on a real repo and real tasks so the Agent becomes a work buddy, not another chat box. Duration and pricing follow team size and scope — book a session first.

Claude Code

Connect day-to-day development to an Agent: understand the repo, change the project, run tests, and finish inspectable work.

  • Repo context and task breakdown
  • Edit, test, and fix failures
  • Permissions, secrets, and safety

OpenAI Codex

Collaborate inside the IDE and the real engineering environment. Completion is not the goal — the Agent has to enter build, debug, and review.

  • IDE and toolchain setup
  • Pairing and code review
  • Work with the tools you already use

AI Workbuddy

Beyond coding-only work: put Agents on everyday tasks — documents, investigations, comparisons, and delivery.

  • From chat to a deliverable
  • Files, tables, and records together
  • A playbook for leads and operators

What every session includes

  • On-site or remote workshop on a real repo or task
  • Clear safety, permission, and data boundaries
  • A keep-using checklist for both leads and engineers

FDE METHOD / 03

The standard FDE engagement

We do not start by building a generic AI platform, and we do not ask the business team to write a complete “AI spec” first. We start from real work, find the highest-value tasks that fit Agents, test them with frontier models, then turn what works into reusable workflows, tools, and product capability.

PREPARE / 04

What the customer prepares

No large data dump. No full requirements document. Phase 1 needs these five things.

1. One pilot scope

One product line; one new, in-flight, or reusable historical project; two or three core roles such as R&D, test / validation, customer support, or delivery. Do not cover the whole company in phase 1.

2. A lead and the people who do the work

One business or engineering lead, plus two to five engineers who already run the workflow. They do not need an AI plan. They should do the work the usual way so we can see the workflow.

3. Real, checkable tasks

5–20 tasks: analyze an incident, review a set of results, debug a customer issue, compare two revisions, find a historical issue, judge whether evidence is complete. Prefer clear inputs, a known answer or an expert who can judge, real human time, and a controllable sensitivity level.

4. Minimum data required

Specs, technical docs, past cases, result files, scripts, tables, and approved internal knowledge. Phase 1 does not need every internal system or the most sensitive data. Provide what the task needs — nothing more.

5. A human baseline

How long it takes today, which steps, which tools, where time goes, where expert judgment is required, and common rework. It does not need to be precise. We can record it together.

HOW WE WORK / 05

Seven stages

STAGE 1 Workflow Discovery

Shadow typical tasks. Capture inputs, what people look at first, tools, judgment points, repeated steps, expert-heavy steps, and the output. Deliverable: Workflow Map + Human Baseline.

STAGE 2 Frontier AI Capability Test

Do not build a complex system yet. Use the strongest current Agents and models, in the allowed environment and data, on the same tasks. Record what was read, which tools ran, what finished alone, where it failed, where humans corrected it, time, and whether the result was acceptable.

STAGE 3 Session & Failure Analysis

Find missing knowledge, repeated human fixes, repeated operations, missing tools, thin context, and unstable judgment steps. The point: why the Agent fails, and what is worth engineering.

STAGE 4 Engineering Enhancement

Turn expert methods into workflows / skills. Wrap repeated work as deterministic tools (parsing, diffs, reports, data handling). Connect a docs repo, Git, database, issue tracker, or business system only when the test shows it is required.

STAGE 5 Golden Task Evaluation

Use 10–30 real cases with known judgment that experts can review. Compare Human Only, Raw Frontier Agent, and Agent + Engineering Harness on success rate, correctness, human interventions, time, evidence, and repeatability.

STAGE 6 Hero Workflow Selection

Keep workflows that have high business value, high human cost, near-usable AI, controllable engineering cost, easy verification, and a path to copy. Stop the rest.

STAGE 7 Productization

Only after real-task proof: an enterprise Agent workspace, fixed workflows, permissions, sessions, evidence, review, and evaluation. The question changes from “can AI finish this?” to “can engineers use it steadily?”

SPRINT / 06

Phase 1 is a 2–3 week sprint

No large program team. The lead picks scope, coordinates engineers, and decides whether to expand. Engineers bring real tasks, work as usual, judge Agent output, and name what is missing. Prefer low-sensitivity, redacted, or historical projects.

Scope

One product line · 2–3 roles · 2–3 workflows · 10–20 real tasks

What you do not prepare

No full AI platform design, no 30-page spec, no company-wide data dump, no “connect every system”, no locked model or final architecture, no large rollout. Prove value first, then fund engineering.

DELIVERABLES / 07

What you leave with

1. Workflow Map

The real workflow and human steps.

2. Human Baseline

Time, steps, and expert dependence.

3. AI Capability Map

What already works, what needs help, what should wait.

4. Failure Analysis

Where the Agent fails and where people correct it.

5. Hero Workflow

The 1–3 workflows worth funding next.

6. Prototype

A running Agent on at least one real task.

7. Evaluation

Human vs AI on real tasks.

8. Phase 2 Roadmap

Systems, data, effort, productization, and ROI.

SUCCESS / 08

Success is not a pretty demo

Phase 1 should answer: which workflows are worth Agents; how far a frontier model gets on its own; the biggest failure modes; how much engineering helps; how much time engineers save; what is worth productizing; and whether phase 2 has a clear ROI.

The loop

Real workflow → frontier test → capture methods → tools and systems → repeatable eval → productize → engineers actually use it → new sessions and feedback → iterate.

The goal

Not a few isolated AI features. A way for the company to keep turning the newest model capability into real productivity.

BOOK A WORKING SESSION

Book training,
or an FDE Discovery Sprint.

Book a meeting