Field Notes
July 27, 20267 min read

The 30-Day Pilot That Actually Produces a Result

Justin Henriksen
Justin Henriksen

Founder & CEO, GetLatest AI

Building an AI agent without the right sequence is like framing a house before pouring the foundation. Looks like progress. Falls over later.

A company a few miles from our office learned this the hard way. Six months. Meaningful budget. Three workflows running at once - lead qualification, follow-up sequencing, pipeline reporting. Week two: data problem kills workflow one. Week six: workflow two breaks on an edge case the demo never showed. Week nine: workflow three sends something to a prospect that was embarrassing enough the rep asked to shut the whole thing down.

Project shelved. Team concluded the technology wasn't ready.

The technology was fine. The sequence was wrong.

They skipped from "this looks promising" to "this runs without us." Didn't verify the data. Didn't test failure cases. Didn't establish what the agent was allowed to do before it did things. Built a production system before they understood the workflow well enough to automate it.

Here's the sequence that works instead.


The 12 steps - and why each one exists

Step 1: Map the actual process. Write down what actually happens - not the official version. The real one. Manual workarounds, system-hopping, the institutional knowledge that lives in one person's head. Those are exactly the things that surface as production failures if you skip them.

Step 2: Identify measurable friction. Find the specific steps where slowness or inconsistency has a number. Average hours between lead submission and first response. Percentage of pipeline deals with no update in 14 days. Follow-up tasks that fell through last month. "It feels slow" is not a baseline. A number you can compare against later is.

Step 3: Select one contained workflow. Not "automate sales." One workflow. Clear trigger. Defined actions. Specific outcome you can check. "When a new inbound lead submits, have a qualified summary and draft response ready for review within 5 minutes." Every additional workflow you add in phase one multiplies your failure surface. Start where friction is highest and the outcome is most measurable.

Step 4: Establish the source of truth. Before any integration work - answer this for every piece of data the workflow touches: which system owns it? Contact exists in your CRM and your email platform with different info? Which is right? No agent resolves conflicting data cleanly. Fix the data before you build the agent.

Step 5: Run the workflow manually with AI assistance. Most businesses skip this step because it feels redundant. It's the most valuable step in the whole sequence.

Don't build an integration yet. Run the workflow by hand. Lead comes in - team member opens it, pastes details into Claude or ChatGPT, gets a scored summary and draft response, reviews it, sends it manually. Do this for every case over 7-14 days.

You find out fast whether the AI output is actually useful. You find edge cases before they break a production system. You get your real baseline - how long the manual version actually takes. Sometimes you find out the workflow isn't worth automating at all because the variation requires human judgment. Better to learn that in week two than after a full build.

Step 6: Define permissions and approval gates. Write down exactly what the agent may do autonomously versus what requires human approval. Be specific. Not "the agent handles follow-ups" - "the agent may draft a follow-up and stage it for review, but may not send any email without explicit approval from the assigned rep." This document is your governance policy. If you can't write it clearly before you build, you're not ready to build.

Step 7: Build the smallest reliable integration. Connect the minimum number of systems needed to make the workflow run. One trigger source. One action destination. One system-of-record update. Not the whole map from step one - just enough to test the pattern. Get the minimal version right first.

Step 8: Test normal and failure cases. Before any volume runs through: test the path the agent is supposed to take, and the paths it shouldn't. What happens when the trigger fires with incomplete data? When the API call fails? When the same lead submits twice? You will hit all of these in the first week of production. A system that handles the happy path but crashes on exceptions is a demo connected to real data.

Step 9: Deploy with human supervision. First week of production is not autonomous operation. A human watches every action, checks every output, logs a verdict for each case: correct, incorrect, or outside scope. Cases where the agent was wrong are not failures - they're the data you need to improve the system. You can't improve what you haven't measured.

Step 10: Measure business outcomes. End of month one - measure the actual outcome you defined at the start. Not "the agent sent X emails." The agent is not the outcome. Speed-to-first-response. Pipeline deal recovery rate. Hours reclaimed for higher-judgment work. If you can't measure it, you either didn't define the outcome clearly or you never established the baseline you needed.

Step 11: Automate additional steps only after reliability is demonstrated. First workflow working - a month of data showing it handles normal cases and escalates exceptions reliably - then expand. Not before. Two weeks of mostly-positive results is not a proven workflow. It hasn't seen enough volume to surface the failure modes that only show up with real traffic.

Step 12: Add adjacent workflows into a coordinated system. As individual workflows prove reliable, connect them. Inbound lead response feeds into qualification. Qualification connects to pipeline monitoring. Pipeline monitoring connects to content distribution. This is how a system grows - not by designing everything upfront, but by proving each component before wiring it to the next.


The 30-day pilot in concrete terms

Week 1: Map the process. Find measurable friction. Select one workflow. Establish your baseline. End of week: you can answer what the workflow is, what success looks like, and how you'll measure it.

Week 2: Run it manually with AI assistance for every case that comes in. Identify data sources. Resolve ownership ambiguities. Write the permissions document. End of week: a working manual version and a clear definition of what the agent may and may not do.

Week 3: Build the minimal integration. Test the normal path and the failure cases you deliberately threw at it. End of week: an integration that handles expected inputs and the inputs you used to try to break it.

Week 4: Go live. Watch everything. Log every case. End of week: your first real sample of production performance data.

At the 30-day mark - one question. Is this working? Not "does the demo still run." Has the workflow run reliably across real cases, and did the business metric you targeted move in the right direction?


Why the sequence compounds

Second agent workflow is easier than the first. You've already established data governance. Already written the permissions framework. Already built monitoring and logging. Already know what failure cases look like and how to test for them.

The infrastructure you built for workflow one reduces the cost of every one that follows.

Businesses that start with one contained workflow, prove it, then expand - they consistently outperform businesses that try to automate everything at once. Not because they move slower overall. They often move faster, because they don't spend months diagnosing which piece of a multi-workflow system broke and why.

One workflow. Make it reliable. Measure it. Then do it again.


One thing to do this week

Pick the single workflow you've been thinking about. Write down: what is the trigger, what is the outcome you would measure, and what is your current baseline for that metric?

If you can't answer all three - that's your next step. Not finding a platform.


Want me to look at yours?

Bring one workflow to a 30-minute call with me. I'll tell you exactly where an agent would fit in your business, what it would take to build, and whether it's even worth it for you. No pitch - just an honest read from someone who does this for a living.

Book a mapping call

Justin Henriksen

Justin Henriksen

Founder & CEO, GetLatest AI

Justin is the founder of GetLatest AI. 25 years building and leading technology, from Principal SWE to CEO. He writes about AI agent architecture, production systems, and what actually works.

Want your AI team built and run for you?

Book a free call30 min · no commitment